Composition and AI
Why AI lays out the median by default, and the layout vocabulary that fixes it.
The silent ignore
Type “design a magazine cover with good composition” into a current image generator. You get a competent-looking cover: title centred in the upper third, image filling the lower two-thirds, subtitle centred at the base. Nothing is wrong with it. Nothing is chosen about it either. The layout is the statistical average of every cover the model was trained on.
Type “design a minimalist app screen” and you get cluttered minimalism: small cards piled at even spacing, every region of the screen occupied, white space implied rather than used. The model treated “minimalist” as a mood word and gave back the median of images labelled that way.
The failure is silent. Type rendering errors are obvious; you can see when a glyph isn’t a letter. Composition failures look fine on first scan. On second look, nothing leads: no focal point, no deliberate tension. The work is the median, and the median has no voice.
It isn’t that AI is bad at composition. It’s that the model defaults to centred, symmetrical, evenly-filled layouts unless you tell it otherwise. As one practitioner analysis of composition in AI imagery put it in February 2026:
“For AI, composition is not automatic. It must be written into the prompt.” (Prashanthi Anand Rao, Medium)
That is this article’s organising idea.
Composition is the fourth Fundamentals vocabulary set, alongside Typography and AI, Colour Theory and AI, and Hierarchy and AI. Same shape every time: name the vocabulary, demonstrate the uplift, apply it. It is the layout-level fundamental, the where it sits layer.
Why composition survives the model
Composition principles don’t vanish in an AI workflow. Christopher Butler, writing on what AI does and doesn’t touch, lists the things that hold: “Structure still communicates before content. Visual hierarchy still guides attention. Negative space still creates rhythm.” The medium changed; the principles didn’t.
They predate the medium by a long way. The Graphic Design and Print Production Fundamentals textbook names balance, hierarchy, and rhythm as compositional principles taught for decades, because they describe what designers do, not what a particular tool allows. Composition vocabulary outlasts model generations.
What you need is a way to hand those principles to the model. The Nielsen Norman Group puts the bottleneck plainly: “Many users may be able to visualize images in their minds but lack the vocabulary to write the required prompt.” Knowing the four composition rules, and the words for them, is what lets you ask for a deliberate layout instead of accepting the average one. Three demos follow. Each names a rule and shows what changes when you do.
A note on the demos: these are practitioner-observed defaults, not controlled studies. The same generator drives both prompts in each demo, so the only variable is the words, not the tool.
Demo 1: magazine cover, rule of thirds
Brief: “Magazine cover for an indie design quarterly, autumn 2026.”
Without the vocabulary. Prompt: “Design a magazine cover with good composition.” Output: title centred upper-third, masthead centred, image filling the lower two-thirds, subtitle centred at the base. Reads as competent. Reads as median. Practitioner write-ups describe exactly this centre-bias: current generators tend to centre the subject and fill the frame unless anchored otherwise.
With the vocabulary. Same generator, same brief. Prompt: “Magazine cover, title at the upper-third focal intersection (the point where the top horizontal third-line crosses a vertical third-line); date strip and issue number anchored to the lower third; generous negative space across the right two-fifths; asymmetric balance carried by the title block.” Output: specific, off-centre, defensible to a client.
Rule of thirds divides the frame with two evenly-spaced horizontal and two vertical lines into nine parts; focal elements go at the four intersections, secondary elements along the lines. Adobe, which teaches the principle to photographers, calls those four intersections the “power points,” and is careful to note it is “not really a rule, more a guideline.” Asymmetric balance distributes visual weight unevenly but in equilibrium: the opposite of centred symmetry.
The reason “upper-third focal intersection” works where “good composition” doesn’t is that one is a location and the other is praise. Praise resolves toward the median of what gets praised, which is centred and filled. A named intersection is a place; the model has something to put there.
Demo 2: minimalist app screen, negative space
Brief: “Minimalist app screen, single-purpose productivity dashboard.”
Without the vocabulary. Prompt: “Design a minimalist app screen.” Output: cluttered minimalism, multiple cards, evenly distributed widgets, every screen region occupied. Practitioner accounts describe it as attempted minimalism: asked for “minimalist,” the output is busy even when the brief isn’t. Minimalism is implied, not deployed.
With the vocabulary. Same generator, same brief. Prompt: “Minimalist app screen: roughly 60% blank canvas; a single focal element in the centre-left; supporting elements at about a third of its scale; generous breathing room between groups; no decorative chrome; let negative space do the structural work.” Output: scannable, single-purpose, recognisably minimalist.
Negative space is the empty area between and around elements, load-bearing in minimalist work, holding the structure that decoration would hold in a busy layout. Proximity, a rule AI defaults rarely honour, is grouping: related elements sit close, unrelated ones are separated. Because the model rarely groups, every widget reads as equally weighted; naming it (“group the input field and its label tightly; separate the action button”) fixes that in one move. (Proximity, white space, and a 4.5:1 minimum contrast ratio for accessibility are the layout fundamentals the Cert IV pathway teaches as a set.)
A specified amount of blank canvas is a constraint the model can act on; “minimalist” alone is a mood it averages. The number isn’t a guarantee of the output, just a target it can aim at instead of guessing.
Demo 3: poster, figure-ground contrast
Brief: “Poster for a contemporary art exhibition opening.”
Without the vocabulary. Prompt: “Design a poster.” Output: a busy, mid-saturation background that competes with the subject for attention. The ground fills with detail rather than receding, so nothing tells the eye where to land.
With the vocabulary. Same generator, same brief. Prompt: “Poster with high figure-ground contrast: a single dominant figure at the upper-left third intersection; a simple, dark, recessive background with minimal detail; strong separation between the figure and what surrounds it.” Output: a clear subject against a quiet ground.
Figure-ground contrast is the visual relationship between the focal element and what surrounds it: high contrast gives a clear foreground, low contrast gives visual noise. Because the model doesn’t separate figure and ground on its own, the load-bearing instruction is the one that names the background as simple, recessive, and minimal, and isolates the single figure. You’re not handing the model a luminance formula; you’re telling it which part of the image is supposed to step back. Naming the relationship is what gets you the separation.
The four rules, as prompt language
What works on a cover works on an app screen and a poster. Composition is one vocabulary set with four rules, and the rules carry across surfaces and generators.
| Rule | What to say in the prompt |
|---|---|
| Rule of thirds | “Place [element] at the upper-third / lower-third focal intersection; anchor secondary elements to the third-lines.” |
| Negative space | “~60% blank canvas; let negative space do the structural work; generous breathing room between groups.” |
| Proximity | “Group [related elements] tightly; separate [unrelated elements].” |
| Figure-ground contrast | “High figure-ground contrast; single dominant figure; simple recessive background, minimal detail.” |
The four rules don’t rank against each other. They solve different layout problems, and no study supports calling one more important than another. What they share is the move that makes each work: replace a praise word with a specific instruction the model can act on.
Composition and hierarchy: the scope split
Composition and hierarchy are sibling fundamentals, and they overlap on a real artefact. The split: composition answers where elements sit (coordinates, ratios, contrast relationships, the layout level); hierarchy answers what’s primary (ranking, scale, attention order, the structure level). A magazine cover uses both: the rule of thirds places the title; the eye-path ranks it against the cover lines. This article scopes to the four composition rules; Hierarchy and AI covers focal points, scale ratios, and reading patterns. Where the demos here lean on hierarchy too, they cross-link rather than redefine.
Where composition sits in the bigger picture
Composition locks into the other three fundamentals. Type lives inside composition: text is a figure, white space is its ground, so figure-ground applies to a wordmark as much as to a poster. Figure-ground contrast is partly a colour decision, since luminance and saturation are colour properties. And composition pairs with hierarchy as described above.
It maps onto the phases in Build Taste, Generate, Refine: in Phase 1 you collect references for their compositional moves; in Phase 2 the vocabulary becomes prompt constraints; in Phase 3 the close pass catches composition drift. AI multiplies the composition you specify: without the vocabulary, evenly-filled centred work, faster; with it, deliberate layout.
TGDS Verdict
Composition is the fundamental AI ignores most quietly, because its failures look like competence. Centred, symmetrical, evenly-filled: that’s the default, and it ships every day looking fine and saying nothing. The fix is not a better model. It’s four rules and the words for them: thirds, negative space, proximity, figure-ground contrast. Name the rule and the model has something to honour. Reach for “good composition” and you get the average of everyone else who did.
AI doesn’t lay out a page. Designers who know composition lay out a page with AI.
Cert IV in Design is where this vocabulary gets built: the layout fundamentals in this article are the entry-level set, taught before AI integration so students can name a layout long before they ask a model to produce one.
Next: Hierarchy and AI → (the ranking layer that pairs with this one). Also useful: Build Taste, Generate, Refine (this vocabulary inside a working AI workflow) and Typography and AI (the cluster mate).
Ready to start your design career?
Study graphic design online, at your own pace, with 1:1 support from our Support Angels. Accredited RTO since 2008.
Explore our courses