I spent hours on App Store screenshots, so I built a tool that makes them A developer built sukuji, a tool that generates five App Store screenshots at 1290x2796 from uploaded app screens, using Gemini to plan the set and write headlines and GPT Image to paint backgrounds. The developer found that printing headlines onto the canvas before image generation, defining a clear exclusion zone around the text, and explicitly stating that the printed headline is the one exception all prevented the image model from overwriting or crowding the type. The tool launched first in Japan with English added later, is free with no account required, and is limited to two runs per day due to cost. The most useful thing I learned building this: if you want an image model to work well with text, don't let it touch the text at all. Print the text first, then let the model paint around it. When my app was ready for the App Store, I thought the hard part was over. Then I got to the screenshots. Deciding what each one should say, finding a device frame, lining everything up, exporting at the exact size Apple wants. None of it was hard on its own, but it took me hours. So I built sukuji. You upload a few screens from your app and get five App Store screenshots back at 1290x2796, each with a headline, your screen in a device frame, and a designed background. Under the hood, Gemini plans the set and writes the headlines, and GPT Image paints the artwork. Getting those two to cooperate taught me a few things. Image models have gotten much better at text, especially in English. But I still didn't want the model writing my headlines. I wanted the same typeface and size across all five images, a font I chose rather than one it picked, and text that doesn't change shape every time I regenerate. And a single wrong letter means redoing the whole image. So the headline was always typeset in code and added afterwards. That created a different problem. The model had no idea where the headline would go, so it happily put a floating object right where the words would land. Telling it to "leave the top of the image empty" helped a little, but not reliably. What worked was flipping the order. Now the headline is printed onto the canvas first, next to the phone, and only then does the model paint. It can see the words, so it works around them instead of guessing. Even with the text printed, objects kept crowding it. My prompt said nothing may touch the letters. The model obeyed: things stopped just short of the words, practically leaning on them. Technically, nothing touched. The fix was to stop describing a line and start describing a space: a clear zone around the text, measured out, with the rule that coming close counts as failing. Once "close" was defined as a failure, the crowding mostly stopped. For the opening image there's no phone, just the character from your app on a plain background. My prompt told the model to replace that plain background entirely with a scene. It did, headline and all. The text I had printed onto the canvas was gone. The fix was one sentence: the printed headline is the one exception, and it stays exactly as it is. All three lessons are the same lesson, really. Image models take your words literally, so the exceptions have to be spelled out. Someone on Reddit pointed out that screenshot tools fail when the headline just restates the icon: "Track your habits" for a habit app. They were right. Now each headline has to come from something specific on that screen, and the model asks itself one question before keeping a line: would this work just as well for a competitor? If yes, it rewrites it. I launched in Japan first and added English later. After two weeks: That cost is why it's limited to two runs a day. It's free for now, with no account needed. I'd love to hear where it falls short, especially if the headlines come out generic for your app.