Design Critique with AI
AI gives generic feedback by default. The structure of the prompt is what makes it specific.
The two-prompt difference
Ask a model to review a design with a single line: review this and give me feedback. The reply comes back generic. Improve the contrast. Tighten the hierarchy. Make the call to action clearer. Every line is true of almost any design ever made, which is the problem. None of it touches the design in front of it.
Now give the model the same draft with a different prompt. Paste the brief first: who the product is for, what the brand sounds like, the accessibility floor, the things that cannot move. Then ask it to act as a critical creative director and find the top three failures by category, with specific corrections. The reply changes shape. The points name elements, cite numbers, reference the audience. Same model, same draft. The variable is the structure of the prompt.
AI is good at the broad first pass and weak at strategic critique. Improve contrast is generic; this hierarchy buries the core message for your healthcare audience is what a trained designer says, and the model only reaches it when the context is supplied. The Zoom-In Method (50% to 99% to 100%) makes AI useful as a critique partner by structuring the conversation: full context first, explicit self-review second, human judgment last. It is the Refine phase of The Taste-First Method, the build-then-generate-then-refine workflow this site teaches, covered here in full.
Surface, strategic, structural
Critique splits into three categories, and each one belongs to a different owner.
Surface. Font sizes, padding consistency, alignment, hierarchy slips, the accessibility minimums. These are the catchable defects, and a model finds most of them when asked to review its own output. Take the free wins before spending judgment on the harder calls.
Strategic. Audience fit, brand voice, cultural read, whether the hierarchy serves the message for this reader. The designer’s ground. A model gives generic feedback at best, because clearer means nothing without knowing what clarity has to do for this audience.
Structural. The failures underneath the output: a model agreeing with a bad direction instead of challenging it, bias carried in from training data, confident wrong copy presented as fact. Checking AI output against AI feedback does not catch these, because the same blind spots run through both passes. When AI Gets It Wrong develops this category in full. It exists, and trained perception, not a second prompt, is what catches it.
Knowing which category a question sits in tells you who answers it. Surface goes to the model; strategic and structural stay with the designer. Mix them and the loop runs in circles: the designer argues audience fit, the model keeps returning surface fixes, nothing resolves.
The Zoom-In Method
The method is a three-pass loop. Each pass has one task, and skipping a pass breaks the one after it.
50%: full context dump. The first pass loads the model. Goal and core features, target audience and tone, the colour palette in specific values, every page and component in scope, reference work the design is reaching toward. A model cannot critique what it has not been told, so with no context every critique is generic.
99%: self-review. The second pass asks the model to turn on its own output and find what is wrong, by category, in a numbered list. The numbering is load-bearing: a model honours numbered constraints far more reliably than an open request, which is the mechanical reason give me feedback returns mush. A reusable shape for this pass, generic enough to copy and specific enough to work:
Act as a critical creative director reviewing this design for [audience]. The brand voice is [voice]. The accessibility floor is [WCAG AA / AAA]. Identify, as a numbered list:
- The top three hierarchy failures, each with a suggested scale-ratio correction.
- The top three type decisions that do not serve this audience.
- Any cultural assumptions in the copy worth flagging.
- Accessibility issues: contrast, type sizing, alt-text quality.
- Anything that would pass a WCAG check but fail an audience-specific reading test.
Fill the brackets and the model has something to push against. Leave them empty and it has nothing, so it reaches for the average again.
100%: polish and judgment. The third pass is the designer’s. Apply the surface fixes the model found, and keep the strategic and structural calls in human hands. The model handles whether the kerning is loose at the top. The designer handles whether this typography should be speaking to this audience at all.
Two prompts, one draft
Take one draft: a fictional B2B SaaS landing page. One model, two prompts.
The generic prompt is the single line, review this design and give feedback, and it returns the generic mush from the top of this article: improve contrast, tighten hierarchy, clarify the call to action. Nothing anchored to the design in front of it.
The structured prompt loads the brief first. Audience is mid-market operations directors. Brand voice is calm and confident, not loud. WCAG AA is mandatory. Then the numbered self-review. The response, illustrative of the shape it takes:
- Hierarchy: the H1 dominates at the top, but the three evenly weighted feature columns compete with the primary call to action. A scale ratio closer to 4 : 1.5 : 1 across headline, action, and feature row would make the action read as primary.
- Type: the monospace dateline reads as an engineering-tool tone and works against a calm, confident voice. A tracked sans in caps would sit better.
- Copy: the hero headline uses American-startup idiom that mid-market operations directors are unlikely to read as their own. Reach for plainer language.
- Accessibility: body copy at 60% opacity fails AA contrast over the gradient in the lower third. Lift it to roughly 80%, or drop the gradient behind the text.
- Passes the check but misses the reader: the hands-on-laptop hero illustration signals a junior tool, not a platform an operations director buys. Replace with something more abstract.
These responses are illustrative, written to show the shape context-loaded prompting produces, not a single live run. The logic holds: a model cannot name what the brief never told it, so the right-hand column became reachable only once the context was supplied. The brief is where that context comes from, which is why Brief Writing with AI feeds straight into this loop.
What self-review catches
In JM’s own practice running this loop, a model catches roughly seven in ten of its own mistakes on the self-review pass: font sizes, padding, hierarchy slips, fixed without being told. This is a working observation from doing the work, not a measured benchmark.
The remaining three in ten stay with the designer, along with everything strategic and structural. Self-review clears the surface defects cheaply, before judgment goes to the calls a model cannot make. It is a reason to run the pass, not a reason to trust the model with what comes after it.
The same loop, found twice
The designer Tom Johnson describes a process that lands on the same principle. He starts from a rough draft the AI builds, then turns on it: I start critiquing it, looking at it through a super critical eye and shoot holes. No constructive feedback, just woah this is terrible, fix X, Y, Z. Elsewhere in the essay he frames it as putting on his creative director hat and redlining. The craft stays his; the model carries the rough draft and the surface revision.
Johnson is not citing TGDS, and TGDS is not citing Johnson. Two practitioners reached the same shape: let the model generate the rough version, critique it hard with the context the model lacks, keep judgment in human hands. The convergence is on the principle, not the steps; the structure is the answer to the underlying problem.
The tools that run the loop
The loop matters more than the tool that runs it, but a few are worth naming by use case.
- The Figma design agent. Native to Figma and aware of your components, tokens, and standards. It leans generative, working inside your design system; it can also summarise stakeholder feedback and surface themes. Reach for it when the design system already lives in Figma.
- Impeccable’s critique command. An open-source skill by Paul Bakaus, a command you add to your AI assistant. Run
/impeccable critiqueand it returns a UX review across hierarchy, clarity, and emotional resonance. The clearest demonstration that vocabulary improves AI critique, and it suits design-system and component work. - Custom critics. Build your own with the brief, brand voice, and accessibility floor baked in. Worth it when you critique many designs against the same constraints.
Whichever tool runs the critique, the 50-99-100 structure is what makes the output useful. The tool is never the loop.
Critique is fundamentals, applied backwards
Every critique question is a fundamentals question asked after the fact, in the vocabulary the Fundamentals pillar teaches. Does the scale ratio hold, is the leading right, does the type voice match the reader is typography. Does the palette serve the hierarchy, does the contrast pass, is a training-data bias showing is colour. Is the focal point earned, does the eye path scan, is the level count manageable is hierarchy. Does the rule of thirds hold, is the negative space deployed or just left over is composition. The critique loop is the fundamentals turned around and pointed at finished work.
That is why the loop needs a trained eye at the end of it. The model runs the surface pass; naming what is wrong strategically, and knowing which fix serves the reader, is the judgment a junior designer is still building. The method treats AI as exactly that: a capable junior on the surface, directed by someone who can see what it cannot.
The verdict
AI multiplies the critique you can structure. With no structure it returns improve contrast, useful for nothing. With structure it returns specific, audience-aware feedback on the surface, leaving the strategic calls where they belong. That structure is what the designer brings, and it is learnable.
A Cert IV is where the judgment layer gets built: the trained perception that decides which of the model’s notes to take and which to overrule. The model handles roughly seven in ten. The designer handles the rest, and the rest is the part that matters.
Next: Brief Writing with AI feeds the context this loop runs on. When AI Gets It Wrong covers the structural failures self-review cannot catch. Research and Ideation covers the tools that load the brief.
Ready to start your design career?
Study graphic design online, at your own pace, with 1:1 support from our Support Angels. Accredited RTO since 2008.
Explore our courses