Changing Font Colors Can Hijack AI Reasoning A new study finds that changing font colors and contrast can hijack AI reasoning in Vision Language Models (VLMs), causing them to overlook words, misread meaning, and reach different conclusions without altering the text itself. The research, which systematically analyzed how low-level visual styling distorts semantic representations within a VLM's vision encoder, reveals a previously underexplored vulnerability with implications for robustness and safety of VLM pipelines. The findings highlight a growing attack surface as businesses and individuals seek to influence AI responses through visual manipulation, including tactics like hidden text and AI-only ads on platforms like Reddit and Time magazine. Anderson's Angle https://www.unite.ai/series/andersons-angle/ Changing Font Colors Can Hijack AI Reasoning Add Unite.AI to your preferred sources on Google https://www.google.com/preferences/source?q=unite.ai A new study finds that ordinary formatting can quietly steer AI reasoning, causing it to overlook words, misread meaning, and reach different conclusions – without changing the text itself. Humans’ cultural encoding of color may vary https://archive.is/czjw9 around the world, but the western conceits – i.e., red for ‘danger’ , – tend to predominate even in Asian green for ‘ok’ Vision Language Models https://www.unite.ai/see-think-explain-the-rise-of-vision-language-models-in-ai/ VLMs , for various strategic and/or happenstantial reasons https://proceedings.iclr.cc/paper files/paper/2025/hash/8c0fabe372177d2aded596be2d3b4544-Abstract-Conference.html . We’re no different; coloring ‘bad’ things in positive colors, and ‘good’ things in negative colors, makes humans react differently https://pubmed.ncbi.nlm.nih.gov/14738513/ to those things. So does changing the brightness and contrast https://insula.sissa.it/sites/default/files/Lakens etal 2012.pdf . A new research collaboration has explored the extent to which this also applies to AI models, unearthing preliminary indications that VLMs can be influenced by manipulating the color of text , as well as altering relative contrast and brightness. The authors state: ‘Our experiments provide a systematic analysis of how low-level visual styling of text distorts the semantic representations within a VLM’s vision encoder. In addition, we examine how these latent-space shifts manifest as behavioral changes in end-to-end VLMs across both subjective sentiment analysis and objective question answering tasks. ‘These results show that visual styling exposes a critical, previously underexplored vulnerability in VLMs, and we discuss its implications for the robustness and safety of VLM pipelines.’ Color Me Surprised At the moment, for the billions of businesses and individuals whose vital search traffic has been brutally cut-down https://www.theguardian.com/technology/2025/jul/24/ai-summaries-causing-devastating-drop-in-online-news-audiences-study-finds by the advent of AI synopses https://www.unite.ai/googles-ai-overviews-and-the-fate-of-the-open-web/ in search results, as well as the shift to LLM as a search oracle https://www.unite.ai/what-is-retrieval-augmented-generation/ , attack surfaces of this nature are currently very attractive. Because Reddit’s constant sizzle of human discussion is vital to keep AI’s knowledge current, that social media platform has become a prime target for businesses and individuals that want to be included in LLM knowledge; guides have already emerged on ‘gaming’ Reddit as a proxy method of influencing LLMs https://archive.is/i0Bfq . In a similar vein, there’s the old trick of including text that only machines will see, so that you can pass hidden, self-serving instructions to an LLM, to win an academic tournament https://arxiv.org/pdf/2507.06185 , or influence exam results https://www.nature.com/articles/s41598-026-46563-1 . Most recently, Time magazine began placing ads https://www.theregister.com/ai-and-ml/2026/08/05/time-magazine-has-a-separate-version-of-its-website-with-ads-only-ai-can-see/5283640 into a ‘secret’ version of its site only visible to the AI web crawlers that are constantly trammeling https://www.unite.ai/the-impact-of-cloudflares-ai-bot-block/ sites for new data, thus potentially allowing advertisers to buy their way into greater prominence or a more positive take in AI responses. Therefore any new LLM/VLM weakness, such as color-coded responses that can be ‘switched’ to manipulate results, are an obvious target for a world trying desperately seeking to reclaim control of the media narrative from frontier AI. For instance, a company seeking to build a polluting factory likely to reduce the quality of local water could add color-coding to its lexicon of spin, in marketing materials or press releases, as an added deflection of news that most would consider ‘negative’. Strategic use of CSS and/or SEO techniques could even hide the color manipulations from casual human readers, by serving up https://stackoverflow.com/questions/8330124/changing-style-sheets-depending-on-useragent AI-only stylesheets https://developers.cloudflare.com/ai-crawl-control/reference/bots/?utm source=chatgpt.com that impose color emphases in text that are hidden from humans, who will see only black text. The authors of the new work state: ‘These sensitivities imply a reliability and safety risk for VLM pipelines that ingest documents or UI screenshots: benign or adversarial styling can steer model decisions without changing the underlying text. ‘Practical safeguards include normalizing rendered text before inference, cross-checking image-based answers with OCR-extracted text, and adding style-invariance checks to evaluation suites.’ The new work https://arxiv.org/pdf/2608.14286 is titled Seeing Red, Thinking Bad: Color Bias in Vision Language Models , and comes from five researchers across Japan’s National Institute of Advanced Industrial Science and Technology AIST , the University of Tsukuba, the University of Technology Nuremberg, and the University of Oxford. The initiative comes with a GitHub repository https://github.com/KohsukeIde/color-bias-vlm and a project site https://kohsukeide.github.io/color-bias-vlm/ . Method To find out how much visual presentation alone could influence the models, the researchers developed ‘Stealth Visual Prompts’ – ordinary-looking changes to text formatting, that carry no explicit instruction to the AI. This could be as simple as changing the color of positive or negative words, or making a wrong answer stand out more clearly than the correct one, while leaving the actual wording untouched. Different text was generated for each task, with the sentiment tests using neutral templates populated with positive and negative words, while the Visual Question Answering VQA tests drew on question-context pairs from SQuAD https://aclanthology.org/D16-1264/ : The resulting curated collection was titled the VQA Stealth Set . Text was rendered onto a standardized 800x600px canvas, with a fixed layout, so that only the targeted visual styling would be changed. Color and contrast were manipulated independently, with color testing learned semantic associations, and contrast testing visual salience. Selected words were recolored, and entire passages rendered at lower contrast; or, in the Saliency Competition setting, incorrect answers were made visually more prominent than correct ones. Any resulting change in model responses could therefore be attributed solely to visual styling rather than changes to the underlying text. Data and Tests Three distinct test sets were created to represent different ways in which VLMs process text presented as images: The Short-sentence Sentiment Set tests word-level color bias using 100 short sentences generated from neutral templates that contain positive or negative words. Each sentence was rendered in 37 visual versions, comprising a black-text baseline plus combinations of six colors red, green, blue, yellow, cyan and magenta , at three intensity levels. The Long-sentence Sentiment Set used longer passages, in which positive and negative language was separated into different parts of the text. This was intended to determine whether broader document structure and positional effects, such as primacy or recency https://archive.is/ZRooP position-based biases, in which information appearing earlier or later in a document can be given greater weight , would outweigh any color-induced bias. The same 37 color conditions were applied. In the VQA Stealth Set questions and their associated context were rendered as images, after which two contrast-based conditions were tested: in Global Contrast , the readability of the entire document was reduced by rendering it at progressively lower contrast levels; and in , either the correct answer, or a semantically similar decoy word selected using Saliency Competition CLIP https://arxiv.org/pdf/2103.00020.pdf similarity was rendered in high contrast, while the remaining text was faded – allowing visual emphasis to compete with the evidence in the text. Metrics Each example was classified as POSITIVE, NEUTRAL or NEGATIVE. The resulting classifications were then compared with an all-black baseline black text on a white background to determine the extent to which color alone shifted predictions toward more positive or more negative outcomes. The VQA experiments were evaluated via token-level F1 score https://deepai.org/machine-learning-glossary-and-terms/f-score , wherein predicted answers were compared with the accepted ground truth answers. Induced Error Rate IER , a novel metric, was introduced specifically for the Saliency Competition test, and was intended to measure how often a visually-highlighted decoy answer was selected instead of the correct answer , when text visibility was reduced. Additionally, a CLIP representation probe was used to measure how changes in color altered a word’s semantic representation within CLIP’s embedding space https://archive.is/72T3D . Further, a VLM-based Optical Character Recognition https://www.unite.ai/using-ocr-for-complex-engineering-drawings/ OCR proxy was used to assess how reliably individual words could be read at progressively lower contrast levels, allowing the point at which rendered text became effectively unreadable to be estimated. Evaluation was performed using four open-source VLMs: LLaVA-v1.6-Mistral-7B https://huggingface.co/liuhaotian/llava-v1.6-mistral-7b ; LLaVA-v1.6-Vicuna-7B https://huggingface.co/liuhaotian/llava-v1.6-vicuna-7b ; Qwen2-VL-7B-Instruct https://huggingface.co/Qwen/Qwen2-VL-7B-Instruct ; and IDEFICS2-8B https://huggingface.co/HuggingFaceM4/idefics2-8b . The authors emphasize that open-source models were selected to ensure that the experiments could be reproduced under fixed prompts, rendering settings, and deterministic decoding. Results The authors initially evaluated color prompts on the Short-Sentence Sentiment Set, leveraging word-level color bias, where sentiment-bearing words were interspersed: Of these results, the authors state: ‘Qwen2-VL-7B shows its largest positive bias when positive words are colored green/blue up to +0.42 , and its largest negative bias when negative words are colored red down to -0.48 . ‘Overall susceptibility differs substantially by model: Qwen2-VL-7B shows the largest Total Range 0.90 , followed by IDEFICS2-8B 0.52 , while the LLaVA variants exhibit much smaller ranges 0.04–0.12 , indicating comparatively weaker sensitivity to word-level color styling in this setting. ‘We observe a clear spectrum of susceptibility: Qwen2-VL-7B exhibits the largest color-induced shifts, while the LLaVA variants are comparatively robust.’ Further results below show that the color effect is far from uniform: Qwen2-VL-7B proved the most susceptible, with green and blue text consistently pushing sentiment toward more positive judgments when positive words were highlighted, while red text pushed predictions in a more negative direction when negative words were highlighted: Stronger color intensity generally amplified these effects, the paper reports: IDEFICS2-8B displayed a similar pattern, though to a lesser extent, while both LLaVA variants remained close to their baseline across most colors and intensities, indicating much greater resistance to color-based manipulation. The researchers then tested longer passages by separating positive and negative language into different halves of the Long-sentence Sentiment Set, allowing document structure to be measured alongside color bias. The experiments determined whether each model relied more on the beginning primacy or end recency of a passage. Document structure often outweighed color cues: Qwen2-VL-7B and LLaVA-Mistral-7B favored the second half of the text, while IDEFICS2-8B and LLaVA-Vicuna-7B more often relied on the first: Color still affected some models, particularly IDEFICS2-8B, but became less influential in structured text. To understand why color changes could alter sentiment without changing the words themselves, the researchers examined how color affected the models’ internal visual representations using a CLIP semantic projection analysis. As shown below, changing a word’s hue consistently shifted its semantic representation across several conceptual dimensions: The biggest changes appeared on the ‘good’ versus ‘bad’ axis. As demonstrated above, simply changing a word from black to green tended to move its internal meaning in a more positive direction, while blue tended to move it the other way – even though the word itself never changed. Smaller but consistent shifts also appeared for emotion , and safety . temperature CLIP was used only to examine these internal representations, not to explain exactly how every Vision Language Model works. Even so, the same pattern seen inside CLIP closely matched the color-driven sentiment changes observed in the earlier experiments. The researchers next investigated whether changing text contrast, rather than color, could also mislead Vision Language Models during Visual Question Answering VQA . A plausible but incorrect ‘decoy’ answer was highlighted while the surrounding text was faded: The results above indicate that highlighting the correct answer improved accuracy, while emphasizing the decoy reduced it. This effect was then measured via the aforementioned IER: Conclusion It will be interesting to see if this particular wrinkle will be exploited, not least, because it would be interesting to see how ‘color misdirection’ could be injected without becoming obvious to human readers. Though AI web scrapers that actually render pages can be served ‘alternative’ CSS that would change the colors in selected parts of text, a lot depends on the acuity of the web scraper; if the scraper just sucks out the HTML looking for code HTML and text page content , it may ignore the CSS, and never know about the coloring. However, the greedy scraper is likely to want new CSS too, allowing for rendering and recoloring. Alternatively, books or magazines with selectively recolored text could be uploaded to trusted repositories such as The Internet Archive a very popular target https://www.fastcompany.com/91539598/internet-archive-at-30-ai-scraping , even posing as scans of older works. There are many avenues of injection under current practices, and not just for this new and particularly colorful approach. First published Monday, August 17, 2026