Plenty of owners hesitate to let AI draft their Google review replies because they assume a customer will spot it and hold it against them. There's real data on this now, and it says the opposite of what most owners expect. When BrightLocal put a real business owner's reply next to an AI-written one, without telling anyone which was which, more people picked the AI version. Twice, in two separate years.
BrightLocal's blind test: what actually happened
BrightLocal runs one of the largest recurring surveys of local review behavior, the annual Local Consumer Review Survey, the same source already behind the finding that half of consumers are put off by a generic reply (see our guide to fixing robotic-sounding replies). In 2024, and again in 2025, the survey added something new: a blind test.
Panelists were shown a real customer review, then two possible replies to it: one written by the actual business owner and pulled straight from Google, the other generated by feeding the same review into ChatGPT with a simple prompt. Nobody was told which reply came from which source. They were just asked which one they'd rather receive.
The 2024 test used a restaurant review. The 2025 test switched industries entirely and used a veterinary clinic. Both times, the result was the same: 58% picked the AI-written reply, 42% picked the human one. Different business, different review, different year, same split.
Why the survey answers and the blind test disagree
This sits oddly next to how people describe their own feelings about AI. Ask consumers directly and plenty say they trust AI-generated content less, or that they'd rather know when something wasn't written by a person. Ask them to pick between two unlabeled replies, and a majority choose the machine. That gap isn't really a contradiction, it's a measurement problem: a survey question asks people to judge a category, 'AI,' loaded with everything they've read about chatbot spam and fake reviews. A blind test asks them to judge one specific piece of writing against another.
The category judgment carries all that baggage. The specific judgment is just: does this reply sound like someone read what I wrote? Most owner replies, it turns out, don't clear that bar as often as owners assume, which is also why BrightLocal's wider 2026 survey found half of consumers are put off by a generic or templated reply, and roughly 89% read a business's replies in the first place. People are paying attention to the wording far more than to the byline.
What actually decides trust: specificity, not authorship
The honest reading of BrightLocal's result isn't AI writes better replies than people do. It's that authorship is nearly undetectable, and specificity is what people are actually grading. A prompted AI reply that references the exact thing a reviewer said will usually beat a rushed, copy-paste human reply to the same review, and it will lose to a human reply that does the same thing.
“Customers can't reliably tell a good AI reply from a good human one. They can tell a bad one from a bad one, every time.”
Same review, two ways to answer it
Here's the difference the blind-test data is actually picking up on. Same review, two replies, and either one could have come from a person or a tool.
“Fixed my bike same day, which I didn't expect, but nobody mentioned the brake pads would cost extra until I was already at the register.”
Thank you for your feedback! We take all customer experiences seriously and are always working to improve.
“Fixed my bike same day, which I didn't expect, but nobody mentioned the brake pads would cost extra until I was already at the register.”
Priya, you're right that the brake-pad cost should've come up before we started the work, not at pickup. I've asked the front desk to flag any add-on cost before it happens from now on. Glad the same-day turnaround worked out, that's the part we try hardest to get right.
Run both replies through a blind test and the second one wins, regardless of which one a person typed and which one a tool drafted. That's the actual lesson in BrightLocal's data: stop worrying about the byline and start checking whether the reply proves someone, or something, actually read the review.
- Reference one concrete detail from the actual review: a name, a number, a specific complaint
- Vary the opening line so replies don't read like the same one pasted repeatedly
- Let low-stakes, everyday reviews post automatically once a reply clears the specificity check
- Hold negative, emotional, or high-stakes reviews for a human read no matter how well the draft tests
- Assume a customer can tell, or would care, whether AI or a person drafted the reply
- Let a reply concede a fact nobody has verified, especially under a contested negative review
- Reuse the same sentence across multiple replies; blind testers notice generic phrasing even without naming it
- Treat 'sounds human' and 'is accurate' as the same requirement, a reply needs both
Do you have to tell customers a reply is AI?
No, and this isn't a gray area. Google's own review-reply policy doesn't ask who or what wrote a reply, it asks whether the reply is accurate and not deceptive. There's no authorship disclosure requirement, no 'written with AI' label to add, nothing extra to check.
A 3-question check before any reply posts
None of this means every reply should post itself without a look. It means the question worth asking isn't 'did AI write this,' it's whether the reply would survive being read next to the review it's answering. Three questions cover it:
- Does it mention something specific from this review, a detail, a name, a number, not just the star rating?
- Could this exact sentence sit under a different review with the name swapped? If yes, it needs a rewrite before it goes near a customer.
- Is this review negative, emotional, or about money, health, or safety? If yes, a human reads it before it posts, no matter how well the draft tests.
The first two questions are about quality, and either a person or a tool can pass or fail them. The third is about risk, and it's the one rule that doesn't bend: sensitive reviews always get a human look before anything goes live, because a wrong-toned reply under a one-star complaint does more damage than a slow one.
Putting this to work on your own reviews
In practice this is a simple split. Everyday reviews, positive or neutral, with no fact at stake, can be drafted and posted automatically as long as each one references the actual review instead of a template. The sensitive minority waits for a one-tap human OK before it goes public. That's the shape of how Resparo works: it drafts every reply from the real text of the review in your voice, holds anything negative or high-stakes for your approval, and lets the everyday ones go out on their own.
You don't need to take our word for whether a drafted reply clears the specificity bar. Paste one of your own reviews into the free reply generator and read the draft back against the three questions above. If it passes, BrightLocal's data says your customers won't care who or what wrote it. If it doesn't, that's worth fixing whether a person or a machine drafted it. For where automation helps and where it doesn't, see our guide to automating review responses, and if you're still weighing whether AI replies are worth it at all, this is the honest breakdown.
