Most generated images come out centered. The subject sits in the middle, framed from roughly chest up, shot from eye level, with everything else arranged politely around it. That is the default because it is the average of everything the model has seen, and unless you push against it, the average is what you get.
Pushing against it is mostly a vocabulary problem. Words that describe where the camera is work well. Words that describe where things sit inside the frame work much less well. Understanding that split saves a lot of wasted attempts.
Camera words move things reliably
Ask for a low angle and the model will actually drop the camera, making the subject loom. Ask for a high angle, an overhead shot or a bird’s-eye view and it will lift. Close-up, medium shot, wide shot and extreme wide shot all change how much of the scene you see and how large the subject sits within it. Over-the-shoulder, profile view, three-quarter view and from behind all reposition the camera relative to the subject, and they land consistently.
These work because they come from photography and film, where they are used constantly in captions and descriptions. The model has seen thousands of images labelled “low angle shot,” so the phrase carries real weight. The same is true of lens language: wide-angle distortion, telephoto compression, shallow depth of field, everything in focus. A wide lens will genuinely stretch the space and make foreground objects bulge.
Distance is worth stating outright rather than implying. “A wide shot of a woman in a field, small in the frame” beats “a woman in a big field,” which usually returns a portrait with some grass behind it. If you want the subject small, say small.
Placement words are unreliable
On the left, in the bottom right corner, in the upper third — these get honored some of the time and produce a mirrored or centered result the rest of the time. The model is not laying out a grid; it is producing something that broadly matches your description, and left-versus-right is one of the weakest signals in it.
If placement genuinely matters, get there indirectly. Describe what occupies the rest of the frame: “a lone figure walking, an empty road stretching away to the right” gives the model a reason to put the figure off-center. Negative space is the ally here — “a single chair against a large empty wall, lots of empty space above it” produces the composition that “chair in the lower left” will not. Give the model something to fill the space with and the placement follows.
Layering is the underused move
Flat, dull compositions usually lack depth cues. Naming a foreground, a middle ground and a background, even briefly, makes an image look composed rather than assembled: blurred leaves in the foreground, a cyclist in the middle distance, city towers behind. Three short clauses do enormous work, because each one gives the model a plane to build on.
The rest is aspect ratio, which many tools treat as a separate setting rather than a prompt word. Composition and frame shape are inseparable — a wide crop invites a horizontal arrangement, a tall one forces the subject to stack. Choose the ratio first and write the prompt to suit it, instead of fighting the frame with words afterwards.
