# DeepSeek Tried to Paint—and Hit a Wall

> Source: <https://www.stork.ai/blog/deepseek-tried-to-paintand-hit-a-wall>
> Published: 2026-10-04 20:21:25+00:00

## A portrait becomes a test of computer use

Matthew Berman gave **[DeepSeek V4.1 Flash](https://www.stork.ai/en/deepseek)** an unusual test: recreate a headshot not by generating pixels, but by operating a web-based replica of Microsoft Paint. This wasn't a standard prompt-to-image diffusion task; the multimodal model had to translate visual input into a sequence of canvas actions.

The challenge demanded more than just image recognition. DeepSeek V4.1 Flash needed to select colors, choose brush sizes, and place strokes on a digital canvas. This experiment probed the model's ability to interpret a visual reference and then execute precise, tool-mediated actions.

Could a fast multimodal model like DeepSeek V4.1 Flash make a coherent image by operating a tool, rather than synthesizing one directly? Berman’s setup aimed to determine if agentic models could move beyond latent space image generation into the realm of **GUI interaction**, using a raster graphics editor to "paint" a portrait.

This test pushed the boundaries of **computer use** for AI, evaluating whether a model could transform visual understanding into an orchestrated series of mouse clicks and brush movements within a virtual environment. It asked if an agent could truly "see" and then "act" to produce a desired visual outcome.

## The result: recognizable, but strikingly abstract

Berman’s verdict on DeepSeek V4.1 Flash’s effort was “actually not bad,” a surprisingly positive assessment given the challenge. The resulting portrait, however, was highly stylized and abstract, capturing an impression more than a likeness. It conveyed broad shapes and a general color palette but lacked the fine details and textural nuance of the reference image.

DeepSeek V4.1 Flash produced a recognizable head shape and overall color impression, yet failed to reproduce finer features like eyes, nose, or mouth with any specificity. The absence of convincing texture made the image feel flat, an abstract interpretation rather than a detailed recreation. This output highlighted a critical limitation in its approach to GUI-driven visual tasks.

Observing the video, the model clearly struggled with the iterative layering technique common in human-driven digital painting. Unlike systems such as Astra, which employs sequential, overlapping brush strokes to build shading and texture, DeepSeek V4.1 Flash generated broad, simplistic forms. This mechanical execution resulted in a portrait that, while identifiable, lacked intricate detail.

The experiment wasn’t a test of DeepSeek’s general image-making capabilities, but rather its ability to translate visual observation into precise, sequential actions within a raster graphics environment. The outcome underscored a gap in agentic motor control and spatial-temporal reasoning, revealing the current ceiling for smaller, faster “Flash” tier models in complex, fine-grained interaction.

## Why layering was the hard part

Placing a single stroke on a digital canvas presents one challenge; composing hundreds of strokes, each building on the last, presents another entirely. DeepSeek V4.1 Flash could identify the general shapes of Berman’s face and lay down broad color fields. What it struggled with was the iterative, layered application of detail.

Berman highlighted this limitation by contrasting it with Astra, Google’s multimodal system. Astra demonstrated a sophisticated understanding of how overlapping brush strokes could create subtle shading and complex textures. DeepSeek V4.1 Flash, by contrast, produced a “very abstract” rendition, lacking the nuanced layering essential for detail.

This isn't just about recognizing a face; it’s about **spatial planning** and **recursive visual feedback**. Successfully painting a portrait with a GUI requires an agent to:

- Precisely control tools
- Track canvas coordinates
- Evaluate each stroke’s impact
- Plan subsequent actions based on that feedback

Without this intricate loop, the agent resorts to simpler, broader gestures. For more on the capabilities and architecture of the models, you can visit the [DeepSeek Official Website](https://www.deepseek.com/). DeepSeek V4.1 Flash, while impressive in its speed and basic tool use, hit a wall where complex, iterative composition was required, revealing a current frontier in agentic motor skills.

Enjoying this? Get one like it in your inbox each morning.

one email a day · unsubscribe in two clicks · no third-party tracking

## The bigger lesson for AI agents

The DeepSeek V4.1 Flash experiment highlights a crucial distinction in AI capabilities. Direct image synthesis, like that from [Midjourney](https://www.stork.ai/en/midjourney) or [Recraft](https://www.stork.ai/en/recraft), generates pixels from a prompt. Agentic action, however, requires an AI to navigate a real interface, making sequential decisions to achieve a goal. DeepSeek had to observe a reference image, then translate that into browser-based commands for a program like **JS Paint**.

This makes the painting task a potent stress test for computer-use systems. Unlike a text answer, the visual output of an agent manipulating a GUI immediately exposes gaps in planning and execution. We can see precisely where the model struggled to translate its understanding into precise, iterative strokes, offering clearer diagnostic feedback than a perfect-looking but internally flawed text response.

DeepSeek’s attempt shows impressive basic tool interaction. It grasped the fundamental task, producing a recognizable if abstract portrait. Yet, the challenge of detailed, iterative work—like layering strokes to build shading and texture—remains a tough bar for current **agentic AI** models. The experiment underscores the complexity of transforming high-level intent into granular, real-world interface actions.

## Frequently Asked Questions

### What did DeepSeek V4.1 Flash do in the portrait test?

It used a browser-based Microsoft Paint replica to create a portrait from a reference image.

### What did the DeepSeek portrait look like?

The result was stylized and abstract, capturing broad shapes and colors but lacking fine detail.

### Why was the portrait less detailed than Astra’s?

The test highlighted DeepSeek’s difficulty planning and layering many precise brush strokes to build shading and texture.

### Is painting with an AI agent the same as text-to-image generation?

No. An agent must operate a drawing interface through sequential actions, while a text-to-image model generates an image directly.
