cd /news/machine-learning/projection-aware-end-to-end-learned-… · home topics machine-learning article
[ARTICLE · art-117296] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Projection-Aware End-to-End Learned Video Compression for 360-Degree Video

A new thesis from arXiv (2608.28689v1) finds that equirectangular and padded equirectangular projections deliver the highest compression efficiency for end-to-end learned 360-degree video compression using the scale-space flow model, outperforming cubemap-based and rhombic dodecahedron projections. The study, which evaluated seven JVET 360Lib projections on JVET test sequences, shows that projection efficiency is codec-dependent, as conventional HM-16.16 codecs favor cubemap-based formats like equi-angular and adjusted cubemap projections.

read1 min views1 publishedSep 1, 2026
arXiv:2608.28689v1 Announce Type: new
Abstract: 360-degree video supports immersive applications such as virtual reality, autonomous driving, and education. Because spherical content cannot be processed directly by conventional video codecs, it must first be mapped to a two-dimensional projection. Projection choice affects spatial continuity, sampling uniformity, motion estimation, and compression efficiency.

This thesis investigates how projection format influences end-to-end neural compression of 360-degree video. Seven formats supported by JVET 360Lib are evaluated using the scale-space flow model, JVET test sequences, and common test conditions. Each sequence is converted from its source equirectangular projection to a coding projection, compressed at multiple rate points, reconstructed, and converted back. Performance is assessed using PSNR, spherical PSNR, weighted spherical PSNR, and Bj{\o}ntegaard delta rate. A differentiable pipeline combining projection conversion, neural compression, and inverse projection is also compared with 360Lib. Results show that equirectangular and padded equirectangular projections provide the highest compression efficiency with the scale-space flow model, while cubemap-based and rhombic dodecahedron projections are less effective. This differs from the conventional HM-16.16 codec, for which cubemap-based formats, particularly equi-angular and adjusted cubemap projections, outperform equirectangular formats. Neural models based on optical flow benefit from the spatial continuity of single-face projections, whereas block-based hybrid codecs better accommodate multi-face layouts. These findings show that projection efficiency is codec-dependent and provide guidance for selecting projections for learning-based 360-degree video compression.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/projection-aware-end…] indexed:0 read:1min 2026-09-01 ·