Cuneiform Analysis with AI
A novel convolution-inspired network structure for classifying ancient cuneiform tablets outperforms the state-of-the-art transformer-based network Point-BERT, according to researchers who developed the architecture. The…
Computer Vision news and analysis on Web Pulse: 3347 curated articles tracking the latest Computer Vision developments, tools, and research, updated continuously from vetted sources.
A novel convolution-inspired network structure for classifying ancient cuneiform tablets outperforms the state-of-the-art transformer-based network Point-BERT, according to researchers who developed the architecture. The…
A user requests an INT8 ConvRot version for FireRed Image Edit v1.1, citing text artifacts and facial drift with standard FP8 variants, and suggests using ComfyUI's native ConvRot support for a quantization pass to impro…
Dukaan Digital has built a prototype that lets kirana store owners digitize their inventory by taking a photo of a wholesale invoice or shelf, using a vision model to extract items and generate an ONDC-compliant catalog.…
GeoAnchor, a new AI framework for 3D reasoning from 2D images, outperforms state-of-the-art models by decomposing spatial information into position, direction, and geometry latents. Developed by an unnamed research team,…
A new benchmark called the Unified Embodied Seeking and Following Benchmark (UESF-Bench) challenges AI to find and follow humans in dynamic settings using only language descriptions, moving beyond current benchmarks that…
Researchers propose JITOMA (Just-In-Time On-demand Memory Activation), a closed-loop framework that unifies task reasoning, perception, and memory to combat perceptual saturation in long-horizon robotics. Instead of buil…
Researchers propose finetuning diffusion generative models with Fréchet Distance loss (FD-loss) to improve synthetic medical image generation for heterogeneous tumors. FD-loss aligns feature statistics of real and genera…
Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family from the Boogu Project, matches or surpasses other open-source models and approaches leading closed-source systems like Nano-Ba…
Researchers propose MGFace, a mask-gated face identification pipeline that predicts mask status and conditionally routes similarity computation, achieving over 80% accuracy with FaceNet and over 90% with ArcFace on the L…
Researchers propose C-Norm (Cell-Distribution Normalization), a method that decouples abnormal and normal cells from ThinPrep Cytologic Test (TCT) images and re-synthesizes them to ensure uniform cell distribution, impro…
Researchers propose Samba, a hybrid Mamba architecture for audio-visual navigation that replaces conventional GRUs with an adaptive selection-enabled Mamba State Encoder (M-SE) and introduces an Audio Mamba Encoder (AME)…
Researchers propose Differentiable Geometry Image (DiffGI), an end-to-end 3D-to-2D mapping framework that replaces binary occupancy maps with a continuous 2D Truncated Signed Distance Function (TSDF) to eliminate stairca…
Researchers propose Temporally Consistent Universal Adversarial Perturbations (TC-UAP), the first protection method against both reference- and tuning-based video customization, addressing three key challenges: image-lev…
Researchers from an unnamed institution have developed ProcessSynthesizer, a method that uses large language models to generate procedural materials by reasoning over expert demonstration traces, outperforming prior grap…
Researchers present AffectFlow-DINO, a multi-task learning system for the 11th ABAW challenge that uses a conditional rectified-flow head to model ambiguity in facial behavior, achieving a multi-task learning performance…
BitMind Forensics (BMF), trained through Bittensor SN34's open adversarial competition, achieves 0.936 AUC on Sumsub's original images and 0.872 pooled AUC over its full manipulation battery, outperforming state-of-the-a…
A Transformer-based Masked Autoencoder achieves 91.3% Hungarian matched accuracy in unsupervised grouping of steel surface defects, according to a new arXiv preprint (2607.13178v1). The method, which masks 75% of input i…
Researchers at MPI Informatics have developed a differentiable polarized path tracing method that enables stable gradient estimation for inverse rendering, overcoming numerical instability caused by rank-deficient polari…
Researchers introduce Self-Correcting Coupled Markov Jump Processes (SC-CMJP), a framework enabling concurrent image understanding and generation through cross-modal coupling, and its training-free sampler CO₂Jump achiev…
A new study from researchers at arXiv introduces the Visual Dependency Gap (VDG) metric to audit whether video large language models (LLMs) rely on visual information for correct answers. Testing twenty models from 2B to…