cd /news/artificial-intelligence/coanerv-coordinate-aware-token-space… · home topics artificial-intelligence article
[ARTICLE · art-99291] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

CoANeRV: Coordinate-Aware Token-Space Neural Video Representation

Researchers propose CoANeRV, a coordinate-aware token-space framework for amortized video representation that forms compact video tokens in one feed-forward pass and uses a shared coordinate-conditioned decoder, avoiding per-video optimization while improving reconstruction quality over prior feed-forward NeRV and INR baselines and reducing peak memory. The code is available at https://github.com/jialong2023/CoANeRV.

read1 min views9 publishedAug 17, 2026

arXiv:2608.13938v1 Announce Type: new Abstract: Neural representations for videos (NeRV) have shown strong reconstruction fidelity by storing video-specific information in network weights. However, existing formulations typically require either costly per-video optimization or video-specific weight generation, making it difficult to scale to efficient amortized video representation. We propose CoANeRV, a coordinate-aware token-space framework that adapts the broader token-conditioned neural-field paradigm to amortized video representation. CoANeRV forms compact video tokens in one feed-forward pass and uses a shared coordinate-conditioned decoder to reconstruct continuous spatio-temporal queries, avoiding per-video decoder optimization or generation while retaining coordinate-level reconstruction flexibility. To make token-space reconstruction effective, CoANeRV introduces a coordinate-aware decoding architecture that aligns spatio-temporal queries with video tokens through axis-adaptive positional encoding and temperature-modulated cross-attention. Block-wise coordinate querying further reduces peak attention memory, making high-resolution reconstruction practical. Experiments on diverse video datasets show that CoANeRV consistently improves reconstruction quality over prior feed-forward NeRV and INR baselines, reduces peak memory compared with attention-based coordinate decoders, and provides efficient amortized encoding without per-video optimization. These results support the proposed video-specific combination of feed-forward token formation, spatio-temporal coordinate retrieval, and memory-bounded dense querying. The code is available at https://github.com/jialong2023/CoANeRV.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @coanerv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/coanerv-coordinate-a…] indexed:0 read:1min 2026-08-17 ·