What is Visual Prompting?
Visual prompting is the deliberate design of visual context to steer a model's attention, constrain its hypothesis space, improve grounding, and shape downstream reasoning, representing a change in the interaction paradi…
Computer Vision news and analysis on Web Pulse: 3346 curated articles tracking the latest Computer Vision developments, tools, and research, updated continuously from vetted sources.
Visual prompting is the deliberate design of visual context to steer a model's attention, constrain its hypothesis space, improve grounding, and shape downstream reasoning, representing a change in the interaction paradi…
GitHub Copilot Vision is generally available as of July 1, 2026, allowing all Copilot subscribers, including free-tier users, to attach images, screenshots, UI mockups, and PDFs into the chat panel for AI-powered code re…
A Chinese research team led by the Institute of Physics under the Chinese Academy of Sciences has developed a fabrication method that cuts production time for 3D optical chips from hours to seconds, as published in the p…
A study accepted at the AGILE 2026 conference found that AI image generators like GPT and DALL-E produce geographically stereotypical and monotonous depictions of places, with newer models showing less diversity than old…
ByteDance's Seeddream 5.0 Pro accepts up to 10 reference images and generates infographics and UI mockups, while OpenAI's GPT Image 2 excels at natural language instruction fidelity and text rendering. The choice between…
NVIDIA released DeepStream 9.1, introducing 13 agentic skills that allow coding agents like Claude Code and Codex to build multi-camera video analytics pipelines from natural-language prompts. The release also adds Multi…
Meta CTO Andrew Bosworth publicly described the company's controversial NameTag facial recognition feature for smart glasses in a podcast interview, weeks after Meta executives claimed the feature did not exist. Bosworth…
Decart AI released Lucy 2.5, an AI model that edits video frames in real time during live streams with sub-40ms latency at 30 FPS. The model uses techniques like MXFP8/NVFP4 quantization and dynamic sparse attention to a…
Ian Macomber released ride-recap, an open-source tool that uses Google's Gemini 3.5 Flash to scan GoPro footage and .fit files from Garmin and Strava, automatically producing a 60-second cycling highlight reel with telem…
DeweyLearn raised $5 million in Series A funding led by SJF Ventures to scale its multimodal AI that grades real-world skills by observing performance against institutional rubrics. The company, named after philosopher J…
PixelUp, a 100% offline AI video upscaler for Windows, uses optimized FSRCNN and ESPCN deep learning models with universal hardware acceleration via CUDA, Vulkan, or CPU to upscale low-resolution videos to high-definitio…
Paris-based Raidium launched its AI-native radiology platform Raidium Read at Moffitt Cancer Center, replacing legacy radiomics tools for clinical trials and research. The platform, built around Raidium's Curia foundatio…
Smart Tech Devs launched ChartAI Pro, an AI-powered realtime chart analyzer on the Google Play Store. The app uses Vision LLMs to turn static chart screenshots into professional-grade technical analysis reports, detectin…
Face AI, a Los Angeles-based face swap platform, announced a major update to its video face swap feature on Friday, improving facial tracking, expression preservation, and scene stability across changing lighting and par…
Apple released the iOS 27 public beta this week, allowing users to install the update immediately rather than waiting for the official September launch. The update features a revamped Siri powered by Apple Foundation Mod…
Fal added Bria AI's Product Dimensions API to its model gallery on July 16, giving developers a hosted endpoint that turns a product photo and supplied measurements into a marketplace-ready image with callout lines, labe…
Google DeepMind reconstructed Pelé's 1959 "lost goal" using Veo 3, Gemini Omni, and Nano Banana Pro, informed by nearly 2,000 historical records and eyewitness interviews. The project, built with Pelé's family, is now on…
Timeline Studio, a local-first AI video editor that runs in the browser, combines a CapCut-style multi-track timeline with browser-side AI voiceovers, automatic captions, vision tools, talking-avatar generation, and dete…
Researchers propose a hybrid variational-generative approach for automated crack detection in digitized paintings, modeling the problem as an inverse image decomposition. The method combines a deep generative prior for t…
Kimi Team released PerceptionBench, a benchmark that isolates atomic visual perception in multimodal large language models by attributing failures across 40+ benchmarks to 10 perceptual capabilities and 3,000 verified qu…