메뉴
HN
Hacker News • 12일 전

JPEG XL 반대론: 웹 이미지 코덱으로서의 재검토

IMP
6/10
핵심 요약

JPEG XL이 기술적으로는 뛰어난 코덱이지만, 웹에서 실질적으로 필요한 용도는 다재다능한 손실 압축이며 JPEG XL의 무손실 압축 이점은 미미하다는 비판적 분석입니다. 최신 AVIF 등 경쟁 인코더들이 지각 최적화와 압축 효율 면에서 JPEG XL을 앞서고 있어, 웹 표준으로 채택할 명확한 근거가 부족하다고 주장합니다.

번역된 본문

웹 이미지 코덱으로서의 JPEG XL의 위치를 조사합니다. 왜일까요? JPEG XL은 기술적으로 인상적인 이미지 코덱입니다. JPEG의 명백한 업그레이드이며, WebP보다 다재다능하고, 웹을 넘어선 사용 사례까지 충분히 대응할 수 있습니다. 하지만 2023년 크롬에서 채택이 거부된 바 있습니다. 로열티 없고, 유연하며, 압축 효율이 높은 JPEG 위원회의 코덱이 대기업들의 주목을 받고도 거부되었기에, 이 결정은 많은 이들에게 납득되지 않았습니다.

최근 Rust로 작성된 JPEG XL 디코더가 어느 정도 형태로 Firefox와 크롬에 들어왔습니다. 웹의 주요 이해관계자들이 새 디코더가 2023년 WebP 취약점 사태의 재발을 막아줄 수 있다는 이유로 JPEG XL에 대한 입장을 되돌리는 것일 수 있습니다. 이것이 웹에서 JPEG XL을 정당화하기에 충분했을까요?

역사적으로 저는 모든 사용 사례에 대해 JPEG XL의 강력한 지지자였습니다. Interop 2024에서 JPEG XL을 지지했고, Jon Sneyers와 Jyrki Alakuijala(포맷의 주요 저자 두 명)와 개인적으로 여러 차례 교류했습니다. 저는 그들의 공개적인 태도, 냉정함, 기술적 역량, 분야에 대한 열정에 지속적으로 감명받아 왔습니다. 이 글은 포맷 저자나 그들의 작업을 폄훼하려는 것이 아니며, 자유 소프트웨어에서 이 코덱이 갖는 상징성과 관련해 어떤 정치적 입장도 주장하려는 것이 아닙니다. 이 글의 취지는 교육적입니다. 2026년 현재 이미지 압축과 웹 플랫폼의 상태를 실증적으로 살펴보고자 합니다. Dmitry Grinberg의 'RISC-V: They Should Have Known Better'에서 일부 영감을 얻었습니다.

웹

저는 이미지 압축 작업을 하고 있으며, 원래는 비디오 압축 분야 출신입니다. AV1 인코더 작업을 하면서 Julio Barba와 함께 AVIF에 상당한 발전을 이루었고 그 과정에서 많은 것을 배웠습니다. 자체 인코더 구축을 시작하기로 했을 때, 어떤 포맷이 가장 높은 상한선을 가지고, 효과적으로 최적화될 수 있으며, 현재와 잠재적 유용성이 가장 큰지 심각하게 고민해야 했습니다. 저는 JPEG XL로 작업하지 않기로 결정했습니다.

볼륨 기준으로, 다재다능한 손실 압축이 대응하지 못하는 웹 사용 사례는 거의 없습니다. 평균적인 웹 사용자는 무손실이 필요 없습니다. 끔찍한 아티팩트(예: 비사진 콘텐츠에 대한 JPEG)를 방지할 만큼 다재다능한 손실 코덱만 있으면 됩니다. 이는 JPEG XL의 무손실 이점을 배제합니다. 실제로 무손실 WebP보다 약 11.9% 작을 뿐이고, 그마저도 웹에는 비현실적인 테스트 데이터셋(157메가픽셀 사진, 10메가픽셀 일러스트, 27메가픽셀 책)에서의 결과입니다. 대역폭 제약에 본질적으로 덜 민감한 사용 사례를 가진 극히 일부 이미지 콘텐츠에서 12%를 절약하기 위해 브라우저에 새 이미지 코덱을 들이는 것이 가치 있다고 할 수 없습니다. JPEG XL은 손실 압축에서 경쟁력이 없으므로, 무손실이 유일한 실질적 이점이 될 수밖에 없기에 이렇게 말하는 것입니다.

손실 압축 효율

JPEG XL의 원래 주장 중 하나는 참조 인코더가 경쟁 인코더보다 지각적으로 더 최적화되었다는 것이었습니다. 이제는 속도와 비트당 충실도 모두에서 다른 인코더들이 더 강력합니다. AV1 참조 인코더는 지각 메트릭에 최적화된 튜닝 모드를 유지하면서 효율을 높이기 위해 통제된 주관적 인간 실험에 기반한 전문 지각 튜닝을 받았습니다. SVT-AV1에도 유사한 튜닝 모드가 있습니다. 최신 인코더가 인간의 눈에 맞게 튜닝되지 않았다는 설득력 있는 주장은 없습니다.

메트릭은 완벽하지 않지만 JPEG XL에게는 험난한 그림을 그립니다. CVVDP, MS-SSIM, SSIMULACRA2 등에서 그렇습니다. aperture-alpha는 Halide Compression의 예정된 인코더로, 코드명 Aperture입니다. libjxl이 최전선에서 경쟁하려면 얼마나 많은 격차를 좁혀야 하는지 보여주기 위해 포함했습니다. 일부 분석은 JPEG XL이 지각적 강점에 비해 메트릭에서 저조하다고 주장하지만, 제가 공유한 그래프 같은 것이 비밀리에 완전히 뒤집힐 수 있을 정도라고 볼 충분한 증거는 보지 못했습니다. CVVDP와 SSIMULACRA2는 매우 강력한 지각 메트릭이며, 차이가 이 정도로 크다면 분명 무언가를 말해줍니다. AVIF의 경우 libaom의 지각 최적화 튠(tune IQ)은 불과 몇 [단계 차이일 뿐입니다].

원문 보기
원문 보기 (영어)
Investigating JPEG XL's place as a Web image codec. Why? JPEG XL is a technically impressive image codec; it is a definitive upgrade over JPEG, more versatile than WebP, and well-equipped to serve use cases beyond the Web. However, it was famously rejected from Chrome in 2023. Because this happened to a royalty-free, flexible, compression-efficient codec from the JPEG Committee that was receiving attention from large companies, the decision didn't land well with many. Recently, a JPEG XL decoder in Rust has made its way into Firefox and Chrome in some capacity. The Web's major stakeholders may therefore be reversing course on JPEG XL given that the new decoder may protect the Web from reliving 2023's WebP vulnerability . Is this all it took to justify JPEG XL for the Web? Historically, I've been a big proponent of JPEG XL for all use cases. I endorsed JPEG XL for Interop 2024, and I've interacted with Jon Sneyers and Jyrki Alakuijala (two of the format's primary authors) personally many times. I'm consistently impressed with their public conduct, level-headedness, technical aptitude, and passion for the field. This piece does not seek to discredit the format's authors or their work, nor to claim any political affiliation relative to the codec's symbolism in free software. The spirit of this post is educational; I want to offer an empirical look at the current state of image compression and the Web platform in 2026. Some inspiration is drawn from RISC-V: They Should Have Known Better by Dmitry Grinberg. The Web I do image compression work , coming from video compression originally. While working on an AV1 encoder , Julio Barba and I made significant advancements to AVIF , and I learned a lot in the process. When I decided to start building my own encoder , I had to think very hard about which formats I felt had the highest ceilings, could be effectively optimized, and had the most present and potential utility. I decided not to work with JPEG XL. By volume, there are very few use cases on the Web that aren't served by versatile lossy compression. The average Web consumer doesn't need lossless; they just need a lossy codec versatile enough to prevent terrible artifacts (e.g. JPEG on non-photographic content). This rules out JPEG XL's lossless advantage, which in practice is only roughly 11.9% smaller than lossless WebP anyway – and on an unrealistic test dataset for the Web (157 MP photos, 10 MP illustrations, and 27 MP books). It cannot be worth bringing a new image codec to browsers to save 12% on a tiny volume of image content with use cases inherently less sensitive to bandwidth constraints. I say this because JPEG XL isn't competitive for lossy, so lossless would be its only real advantage. Lossy Compression Efficiency One of the original arguments for JPEG XL was that its reference encoder was more perceptually optimized than competing encoders. Now, on both speed and fidelity per bit, other encoders are stronger. The AV1 reference encoder received specialized perceptual tuning based on controlled subjective human trials to strengthen its efficiency while maintaining a tuning mode optimized for perceptual metrics. SVT-AV1 has similar tuning modes. There is no compelling argument that modern encoders aren't tuned for the human eye. Metrics aren't perfect, but they paint a daunting picture for JPEG XL: CVVDP MS-SSIM SSIMULACRA2 aperture-alpha is Halide Compression's upcoming encoder, codenamed Aperture. I included it to show just how much ground libjxl needs to make up to compete at the frontier. Some analysis claims that JPEG XL underperforms in metrics relative to its perceptual strength, but I don't see sufficient evidence that this is to the degree that graphs like the ones I shared could be secretly completely reversed. CVVDP and SSIMULACRA2 are very strong perceptual metrics, and definitely tell us something when the differences are this great. For AVIF, libaom's perceptually optimized tune (tune IQ) is only a couple of points lower than its perceptual-metric-optimized tune (tune SSIMULACRA2). Plus, the JPEG XL reference encoder has historically suffered from percep tual issues that remain largely unresolved. There's no such thing as a codec benchmark, only an encoder benchmark; in theory, the ceiling for JPEG XL as a format is higher than libjxl is getting. But how hard would it be to close the gap? As a compression engineer, I believe it is disadvantaged here. Some reasons: JPEG XL doesn't have directional prediction modes. Compressed images are divided into VarDCT blocks (from 2x2 up to 256x256) and transformed into frequency representations of their pixels. Other block-based image codecs like WebP let you predict a block's pixels using surrounding data, subtract this prediction from the actual pixels, and then do the frequency transform. Directional prediction modes can result in blur if your encoder isn't perceptually optimized, but strong mode-decision pipelines can pick the right mode for the job and save lots of bits. For example, edge preservation is stronger in codecs with directional pred, while JXL is weaker here. The proposed solution for the edge-preservation gap is splines, which are vastly more difficult to use. The hard part is on the encoder side: you need an efficient algorithm to figure out which pixels can even be represented as a spline, then feed every candidate through RDO to decide whether it's worth coding. There's no existing PoC for using splines for edge preservation, and I have no reason to believe they'd be better than dir-pred anyway. JPEG XL doesn't have deblocking loop filtering (DLF), or any deblocking filter. It does have two in-loop tools that are sometimes offered as partial equivalents: gaborish, which is the closest thing JXL has to AV1's loop restoration filtering, and EPF (edge-preserving filter), whose closest analogue is AV1's CDEF. Neither is a deblocking filter, and the two together can't fully replace proper DLF. The DLF can smooth images out, but if your encoder is smart it will only help you avoid mosquito noise, which JPEG XL still suffers from. JPEG XL's perceptual "XYB" colorspace is based on a lot of intuition, and doesn't always translate to gains in other formats (like JPEG) even when metrics like SSIMULACRA2 work in the exact same colorspace. The claimed efficiency savings from using XYB also aren't as big as originally advertised because libjxl currently relies on aggressively quantizing the B channel. This has resulted in subpar color preservation, which new JXL encoder developers must explicitly undo. JXL does poorly with non-photographic images. The proposed solution is using patches, but they are more difficult to use than AV1's Intra Block Copy. To get a similar range of expressiveness to IntraBC, the encoder has to deal with additional concepts like layers and blending, which aren't cheap to represent at the bitstream level. Residual coding is awkward. With AV1, you predict a block, subtract the prediction from the source, and the transform coefficients naturally represent the residual. With JXL's construction, you decode a residual frame and then blend a reference patch, so you need an actual frame or layer whose decoded pixels represent the residual. That would likely be a Modular frame, which is interesting because Modular isn't restricted to conventional unsigned image values the way the final rendered image is. An IntraBC block essentially costs a motion vector plus residual coefficients, whereas a JXL construction potentially costs a reference frame, a frame header, a crop, blend information, a patch dictionary entry, patch coordinates, and a residual frame. That overhead can overwhelm the savings unless the repeated region is fairly large or reused many times. Patches have to be explicitly enabled in libjxl below effort 7 because they currently have performance issues. For non-photographic images, the argument that “they should be vector images” doesn't hold up because ma