Text inside generated images used to be a running joke — signs covered in letter-shaped scribble, book covers written in a language that does not exist. Newer models handle words far better than they used to, but text is still the most fragile thing you can ask for, and the difference between a clean result and a mangled one is largely in how the request is written.

Keep it short and quote it

The single strongest predictor of correct spelling is length. A word or two comes out right most of the time. A short phrase is usually fine. A full sentence starts losing letters, and a paragraph turns into texture. If your image needs a lot of copy, the honest answer is to generate the image without it and add the type afterwards in any editor — that is what most people producing finished work actually do.

Put the exact string in quotation marks and say where it goes: a sign that reads “OPEN”, a mug with “Monday” printed on it. Quoting separates the words you want rendered from the words describing the scene, which is exactly the confusion that produces images where your caption ends up illustrated instead of written.

Say it once. Repeating the string, spelling it out letter by letter, or restating it at the end of the prompt tends to produce duplicates — the same word appearing twice on the same sign, or ghosted behind itself.

Give the text a surface

Words render better when they have somewhere obvious to live. A sign, a label, a banner, a book cover, a screen, a chalkboard, a t-shirt — each is a flat, framed surface the model has seen carrying text thousands of times. Floating words with no home are much less stable, and text wrapping around a curved or angled object is the least stable of all.

The surface also lets you control the look without fighting the letters. Bold sans-serif lettering, hand-painted script, neon tube sign, stencilled block capitals, embroidered patch — style words like these steer the typography, and asking for a heavier, simpler letterform generally raises the odds of correct spelling. Thin decorative scripts and ornate serifs break down first.

Size matters in the same way. Text that occupies a decent share of the frame comes out cleaner than text rendered small in the background. If the words are the point of the image, frame the image around them.

Everything else should be silent

Most garbled text in an image is text nobody asked for. A street scene generates shopfronts, a desk scene generates paperwork, a laptop generates an interface — all of it filled with invented lettering. When you only want one readable string, it helps to say the rest carries no text: blank signage, unmarked packaging, a plain screen.

And check what came back before you use it. Models are confident about spelling in a way that does not correlate with being right, and a transposed letter in a short word is easy to miss on a first glance. Read the words in the image out loud once; it catches most of what the eye slides past.

keyboard_arrow_up