about
Hybrid Video Compression Framework Using Reference-Guided Restoration Network (arxiv.org)
2 points by ksec on Apr 15, 2023 | hide | past | pdf | 2 comments on HN

In plain words: The system sends a compressed video plus one frame saved without quality loss, then a two-step cleanup network uses that sharp frame to rebuild lost details. It matched the best deep-learning codecs in quality while running faster and lighter, and plugs into standard encoders.

Abstract · Lightweight Hybrid Video Compression Framework Using Reference-Guided Restoration Network

Recent deep-learning-based video compression methods brought coding gains over conventional codecs such as AVC and HEVC. However, learning-based codecs generally require considerable computation time and model complexity. In this paper, we propose a new lightweight hybrid video codec consisting of a conventional video codec(HEVC / VVC), a lossless image codec, and our new restoration network. Precisely, our encoder consists of the conventional video encoder and a lossless image encoder, transmitting a lossy-compressed video bitstream along with a losslessly-compressed reference frame. The decoder is constructed with corresponding video/image decoders and a new restoration network, which enhances the compressed video in two-step processes. In the first step, a network trained with a large video dataset restores the details lost by the conventional encoder. Then, we further boost the video quality with the guidance of a reference image, which is a losslessly compressed video frame. The reference image provides video-specific information, which can be utilized to better restore the details of a compressed video. Experimental results show that the proposed method achieves comparable performance to top-tier methods, even when applied to HEVC. Nevertheless, our method has lower complexity, a faster run time, and can be easily integrated into existing conventional codecs.

Hochang Rhee, Seyun Kim, Nam Ik Cho
arXiv:2303.11592 · eess.IV, cs.CV · submitted Mar 21, 2023
abstract · pdf · html

add comment on HN

Oh, and another point: instead of bundling an image with the video stream, they could "splice" in near lossless I-frames into the video stream wherever there is a scene change, then extract them on the decoder end. This would have some significant advantages:

- Full backwards compatibility. The "AI decoder" encoded video would play in any regular video player

- No wasted bitrate encoding the same frames twice.

- No complicated colorspace conversion trickery needed to match the (RGB?) JXL image with the video.

- Theoretically, a full GPU workflow with no CPU needed, though researchers would probably do most work on the CPU if prototyping it with VapourSynth or PySceneDetect.

- The DL network could be formatted/trained in YUV420 (or a colorspace with a similar shape) instead of RGB, which could make it much faster/smaller.

This is not that crazy either. av1an already does something like this in AV1 and HEVC, splitting the video into scenes, running the jobs in parallel and then splicing them together.

> We use PySceneDetect to iden- tify dynamic videos with a scene change and supply a new reference frame at the point of the scene change.

> We adopt PSNR and MS-SSIM [44] for evaluating video quality

This is a pet peeve of mine. The standard for the video ML community seems to be PSNR/MS-SSIM, PNG dumps and "basic" tools like pyscenedetect, but I have yet to see a single paper use something more performant and flexible like vapoursynth, and only a few seem to use butteraugli, vmaf, or even just newer SSIM variants.