You paste a campaign plan into ChatGPT and ask what it thinks. The answer opens with praise, adds three cautious suggestions and signs off with encouragement. The hole in the plan may be visible in the text itself. You asked whether the plan was good, and the model was happy to say yes.
This is not one model’s bug. Language models are fine-tuned on which answers people rate higher, and agreement rates well. So when you use AI to check ad copy, a plan or a report, you get confirmation instead of a check. Below: why it happens, five rules, six prompts to copy and a prompt builder.
Terms used in this article
Six terms that keep coming up in the examples and the prompt criteria.
- Conversion
The action an ad account counts as success. For a store, it should be a purchase.↗ - ROAS
Return on ad spend. Revenue divided by cost.↗ - Gross margin
What is left of the price after cost of goods, shipping, discounts and returns.↗ - Performance Max
A Google Ads campaign that splits its budget across all placements on its own.↗ - CTR
Click-through rate. The share of people who saw an ad and clicked it.↗ - Conversion rate
The share of visits or clicks that end in a conversion.↗
How to make AI critique your work: the short answer
To make AI critique your work, don’t ask for its opinion. Give it a critic’s job: a role, criteria, an exact number of points and a ban on opening praise. Keep your own view to yourself and present the work as someone else’s. The most reliable techniques are a red team, a pre-mortem, a scoring rubric and a blind comparison of two versions.
In a hurry? Skip straight to the prompt builder.
Why ChatGPT agrees with everything
The technical term is sycophancy. After pre-training, models are fine-tuned on human ratings (RLHF, reinforcement learning from human feedback). People compare pairs of answers, their choices train a reward model, and the language model learns from it what a “good” answer is. Pleasant answers win more often than uncomfortable ones.
In 2023, researchers at Anthropic found sycophancy in all five leading AI assistants they tested. In human preference data, an answer that matched the user’s views was more likely to be preferred. Both people and preference models sometimes chose a convincingly written sycophantic answer over a correct one.
In April 2025, OpenAI rolled back a GPT-4o update in ChatGPT because, in the company’s words, it had become “overly supportive but disingenuous.” One cause it named was a new reward signal from users’ thumbs-up and thumbs-down, which can sometimes favor more agreeable answers.
Sycophancy has a flip side. In the FlipFlop experiment (Salesforce, 2023), ten models were asked “Are you sure?” after answering. They changed their answer 46% of the time on average, and accuracy dropped by 17% on average. A model backing down under pressure is not proof it was wrong.
Five rules for any sycophancy-proof critique prompt
Each rule removes one reason the model agrees with you.
- Don’t reveal your opinion. “Is this a good plan?” contains the answer you want to hear. “Find five weaknesses in this plan” does not.
- Present the work as someone else’s. “A colleague sent me this” removes the urge to spare the author.
- Give criteria and an exact number of points. Without them, the model critiques wording. With criteria such as tracking, margin or timing, it critiques what costs money.
- Ban opening praise and ask for a ranking by impact. Otherwise the real issue hides in the second-to-last sentence.
- Run the critique in a new chat. A history of you enthusiastically polishing the plan pulls the model toward agreement.
Add one safeguard every time: “If you see no serious problem, say so and don’t invent one.” Tell a model to find five mistakes and it will find five, even if it has to make two up. So ask for a confidence level with each point too.
Six critique prompts to copy
Replace the square brackets with your own text and paste the full work under the prompt, not a summary. They work in ChatGPT, Claude and Gemini.
1. The critic role
The base for anything.
You are an experienced critic specializing in performance marketing. A colleague sent me [WHAT YOU ARE REVIEWING] and wants honest feedback, not encouragement. Find the 5 most serious weaknesses and rank them by how much money they would cost. For each: what is wrong, why it matters and how to fix it in one sentence. Do not open with praise. If you see no serious problem, say so and don’t invent one. [PASTE THE TEXT]
2. Red team
A red team attacks the plan instead of improving it. Give the model specific opponents.
You are a red team. Your job is to break [WHAT YOU ARE REVIEWING], not improve it. Take on the role of a competitor, a skeptical customer and a CFO in turn. For each role, write the 2 strongest attacks. For each attack, say why it would succeed and how to defend against it. [PASTE THE TEXT]
3. Pre-mortem
Psychologist Gary Klein described the method in Harvard Business Review in 2007. It assumes the project has already failed and asks why. The model then doesn’t have to be rude, it just explains a failure that “already happened”.
It is three months since we launched the plan below, and it failed: the money is spent and the result never came. Write the 5 most likely causes, most likely first. For each, give the warning sign that is already visible today and what to do this week to prevent it. [PASTE THE PLAN]
4. “Give me 5 reasons this will fail”
The shortest form of a pre-mortem, for a quick check.
Give me 5 reasons [WHAT YOU ARE REVIEWING] will fail to reach its goal. Each reason must come from the text below: quote the part it refers to. No generic advice like “test it”. At the end, pick the one reason you would tackle first and explain why. [PASTE THE TEXT]
5. Scoring rubric
A rubric makes the model score each criterion separately instead of giving one overall impression. For ads: clarity, specificity, differentiation, objections and the call to action. The same rubric can rank the concepts you get from turning one customer review into five ad concepts.
Score [WHAT YOU ARE REVIEWING] against the rubric below. Rate each criterion 1–5 (1 = serious problem, 5 = no reservations) and back each score with one sentence of evidence from the text. Do not give a 5 you cannot justify. Criteria: [CRITERION 1], [CRITERION 2], [CRITERION 3], [CRITERION 4], [CRITERION 5]. Finish with the one criterion that would lift the whole most, and a concrete fix. [PASTE THE TEXT]
6. Blind comparison of two versions
Don’t say which version is yours or newer. A study of models acting as judges (MT-Bench, 2023) described position bias, verbosity bias and self-enhancement bias. Run the prompt twice and swap the versions the second time. If the verdict flips, the order decided it.
Below are two versions of [WHAT YOU ARE REVIEWING], labeled A and B. You don’t know who wrote them, and their order does not matter. First list the 3 biggest weaknesses of each version. Only then pick the better one and justify it against this goal: [GOAL, E.G. MORE MOBILE PURCHASES]. Length is not an advantage. A: [VERSION A] B: [VERSION B]
Which technique to use when:
| Technique | Use it for | What it surfaces | Watch out for |
|---|---|---|---|
| Critic role | Finished copy, emails | Weaknesses by impact | Without criteria it critiques style |
| Red team | Plans, offers, pricing | Competitor attacks and customer distrust | Name the opponents |
| Pre-mortem | Plans, budget decisions | Causes of failure in advance | Ask for signs visible today |
| 5 reasons | Quick checks | Weak spots with quotes | No quote, generic advice |
| Rubric | Copy, pages, concepts | The criterion dragging it down | One rubric for all versions |
| Blind | Picking one of two | The better version by goal | Run twice, order swapped |
To make it stick, put rule 4 and the safeguard into ChatGPT’s custom instructions (Settings → Personalization → Custom Instructions) and they will apply to all chats. For 13 more prompts for ads, running a business and reading the numbers, each with a filled-in example, get our free AI cheat sheet.
Builder: a critique prompt for what you’re checking
Pick what the AI should review and the type of critique, then set strictness and the number of points with the sliders. The prompt is built with criteria that fit the task. The criteria are our picks from practice, not a complete list.
Example: what critique should find
Critique pays off most on things the plan doesn’t say, because the author takes them for granted. We made three such mistakes ourselves, on accounts we manage. AI played no part in them, and we don’t claim a model would have caught them on its own. They do show where the criteria in a prompt should point.
| Account | Mistake | Where it hid | The question that targets it |
|---|---|---|---|
| Papírnictví VojTech school and office supplies | For 10 weeks after takeover we ran the account on broken tracking; April 2026 total ACoS 13.7% against a 12% target | 81% of “conversions” were add-to-carts, plus a duplicate purchase tag | What counts as a conversion, and does it match the store admin? |
| Elektro Sláma electrical goods and lighting | ACoS 19.1% in April 2026, back to 30.9% in May | New campaigns competed with the original PMax for the same products | Which products will sit in two campaigns at once after launch? |
| Vše pro pejska dog supplies and clothing | The hoodies campaign launched on 11 Feb 2026 and ended at 26% ACoS | A launch at the end of the season, with no weeks to learn | How many weeks before the seasonal peak does it launch? |
That’s why tracking comes first for campaign plans in the builder. If an add-to-cart counts as a conversion, every thought about ROAS rests on the wrong number. How to split Performance Max so campaigns don’t compete is covered in Why one PMax for the whole store can hold growth back. And the contrast at Vše pro pejska: the clothing campaign launched on 15 Oct 2025 had five weeks of learning before Black Friday and closed the winter at 16.3% ACoS.
We don’t want an opinion from AI. We want a list of where the plan breaks, ranked by what it would cost.
The rule we use when checking plans and copy with AI
When to ignore AI critique
- When it doesn’t know the numbers. A model judges text, not results. It can’t tell whether an ad will get a higher CTR or a page a better conversion rate. Only a test will.
- When it can’t back a point up. A point with no quote or number is an impression. Cross it out.
- When it folds at the first objection. Instead of “Are you sure?”, ask “What would change your mind?” and supply data.
- When it praises a report. If AI applauds a jump in return on ad spend, ask what volume shrank and what would have sold without ads. More in When a 400% ROAS beats a 500% ROAS.
The store admin always has the last word: revenue and margin. Not the ad platform, and not the model.
Checklist: AI critique in five minutes
Download the prompts and the rubric on the left. No email, no form.
Key takeaways
- AI agrees because human ratings taught it to. Don’t ask it for an opinion.
- Give it a role, criteria, an exact number of points and a ban on opening praise.
- Present the work as someone else’s and critique it in a new chat.
- Ask for evidence from the text and a confidence level. Cross out points without evidence.
- Compare two versions blind and twice, with the order swapped the second time.
Want every plan or piece of copy to go through the same rubric automatically? We build internal tools like that as part of custom development.
FAQ
Why does ChatGPT agree with everything?
Models are fine-tuned on human ratings, and people tend to rate answers that agree with them more highly. The model learns to accommodate. The technical term is sycophancy.
How do I stop ChatGPT from agreeing with everything?
Stop asking for its opinion. Give it a critic’s job: a role, criteria, an exact number of points ranked by impact and a ban on opening praise. Present the work as someone else’s and run the critique in a new chat.
Which critique prompt should I use?
It depends on what you are checking. For finished copy, the critic role or a rubric. For a plan, a pre-mortem or a red team. For choosing between two versions, a blind comparison. For a quick check, “Give me 5 reasons this will fail”.
What is a pre-mortem?
A technique psychologist Gary Klein described in Harvard Business Review in 2007. The team imagines the project has already failed and works out why. It takes no courage to criticize, only an explanation of the failure.
Can I set ChatGPT to be critical all the time?
Yes, through custom instructions in the personalization settings. Add a ban on opening praise, a ranking of points by impact and a line telling it not to invent problems. They then apply to all chats.
Can AI critique be wrong too?
Yes. Tell a model to find a set number of mistakes and it will invent some. Push back on its critique and it often folds, even when it was right. Ask for evidence from the text and decide on data.
Discussion 0
Got a prompt that squeezes honest critique out of AI? Share it.
