Show HN: Stemning AI
Stemning AI launched a new product called Historical Portraits that transforms user photos into six decade-specific styles from the 1950s to the 2000s, offering one-click generation and 4K downloads with no subscription …
Computer Vision news and analysis on Web Pulse: 3349 curated articles tracking the latest Computer Vision developments, tools, and research, updated continuously from vetted sources.
Stemning AI launched a new product called Historical Portraits that transforms user photos into six decade-specific styles from the 1950s to the 2000s, offering one-click generation and 4K downloads with no subscription …
Around 77% of vision AI implementations in manufacturing never make it past the pilot phase, according to Roboflow, citing a failure of integration rather than technology. Jeff Witt, Digital Transformation Leader at a Fo…
LogoTeddy launches an AI-powered logo animation tool that generates custom cinematic reveals from static logos without templates. The service, which costs one credit ($0.01) per 5-second 480p reveal, offers free 100 cred…
Dutch construction robotics startup Monumental has raised a $32m Series B led by Khosla Ventures to deploy more of its bricklaying robots on building sites in Britain and, for the first time, in the United States. The Am…
A new category of wearable devices combining open-ear headphones with built-in cameras is emerging, with products like Auriview's X1 and Rollme's Aircam offering AI-powered computer vision and audio playback. These gadge…
Google released GNM Head, a high-fidelity statistical 3D model of the human head, as the first open-source component of its GNM Ecosystem. The model provides fine-grained, disentangled control over identity, expressions,…
RF-DETR is the strongest starting point for most computer vision projects in 2026, topping both COCO and the real-world RF100-VL benchmark for object detection and segmentation, with SAM 3, DINOv3, GLM-OCR, and Gemini 3.…
Hugging Face's top AI papers for July 15, 2026, highlight trends in long-horizon agents, robotics foundation models, video understanding, and efficient training. Notable papers include Direct On-Policy Distillation (Dire…
Google Photos' "Ask Photos" feature will soon let users edit photo metadata—including timestamp, caption, and location—by simply asking, according to an APK teardown of version 7.84.0.947289513 by Android Authority. The …
Pixel Rehab, a free browser tool, recovers true pixel art from AI-generated images by detecting the hidden grid and rebuilding the artwork at its native resolution, running entirely locally without uploading any data. Th…
Object detection models have evolved from two-stage detectors like R-CNN to one-stage models such as YOLO and transformer-based architectures like DETR, each offering trade-offs between accuracy and speed. Anchor-free an…
FFAvatar introduces a Transformer-based 3D Gaussian framework for constructing animatable 4D head avatars from a few reference images, with quality improving as more images are added. The system uses an alternating atten…
RESOURCE2SKILL, a new framework from an unnamed research team, transforms tutorial videos and multimodal content into executable skills for AI software agents, boosting agent performance by an average of 11.9 percentage …
A new benchmark called FlipSet reveals that most vision-language models fail at Level-2 visual perspective taking, with 75% of errors stemming from egocentric bias. In an evaluation of 103 models, the majority performed …
Researchers have developed GIAVA (Gaze Integrated Active-Vision ALOHA), a robotic vision system that mimics human foveated vision by integrating eye-tracking and perspective control data from human operators into Vision …
EzTranslate, a free browser-based photo translation tool optimized for Traditional Chinese (Taiwan) speakers, launched on Show HN. The service uses AI to recognize and translate text from images, supporting English, Japa…
A new framework called IQA-T1 challenges traditional image quality assessments by integrating explicit perceptual observations with multimodal reasoning, achieving the best overall performance across seven IQA benchmarks…
A new study reveals that intermediate representations from models like CLIP and DINOv2, combined with semantically aligned references, significantly improve synthetic image attribution accuracy. The research shows that a…
A developer released PixFinder, a free Windows tool that combines AI (SigLip2) and OCR to search local images by content rather than filename, running entirely offline with no cloud or tracking. The tool works smoothly f…
South Korean AI startup ModigenceVision won the Korea-Germany Connect: AI Startup Pitching Challenge 2026 on Wednesday, earning the right to present its 3D vision systems for robotics and industrial automation before abo…