A ChatGPT prompt can rank faces roughly the way people do, but it cannot rate them the way people do, and it does not measure anything. That is the finding of QOVES's own research page, Do AI models rate faces like humans? (labelled 2026 / Summer, read 17 September 2026): "AI models do not judge facial attractiveness like humans; they rate almost everyone above average and never give the lowest scores. However, models can rank faces similarly to how people do." On a 1–7 scale the human average was 3.02; ChatGPT averaged 4.41, Claude 4.53, Gemini 4.69 and Grok 5.21.
So the honest answer to "QOVES vs ChatGPT" is that they are different kinds of thing. A prompt gives you a fast, flattering opinion. QOVES sells a measured report reviewed by people, for $150 a year and a wait of up to 28 days. This page sets out what the data says, what a prompt structurally cannot do, and which option fits which need.
Part of the SoftMaxx facial anatomy glossary — every term, one definition each.
1. What QOVES tested, in their words
The study used the Face Research Lab London Set: "102 neutral, front-facing portraits normalized for pose, lighting, and framing. 2,513 people rated each on a 1–7 attractiveness scale." Then "four commercial models (Claude, ChatGPT, Gemini, and Grok) rated the same faces one at a time, with no memory of the faces they had already seen. Each model rated every face repeatedly," and the repeated ratings were averaged into one score per face. (qoves.com/research, read 2026-09-17.)
| Rater | Mean rating (1–7) |
|---|---|
| Humans (2,513 raters) | 3.02 |
| ChatGPT | 4.41 |
| Claude | 4.53 |
| Gemini | 4.69 |
| Grok | 5.21 |
Two more lines from the same page matter: "No AI model rated anyone a 1 out of 7, and only Grok gave the highest rating." The models compress everyone into the upper middle of the scale.
2. Ranking is not rating
The same page reports rank agreement between each model and the average human of Spearman ρ between .58 and .76 (Claude: ρ = .76, 95% interval .67–.83), and comments: "Rank agreements between .58 and .76 are not trivial. All models, except Grok, order the faces in terms of attractiveness close to how humans do."
Read together, the two results say something specific. Ask a chatbot which of two photos of you is stronger and the answer carries real signal. Ask it for a score out of 10 and the number is inflated by roughly one and a half to two points on a seven-point scale, with the bottom of the scale never used. A rating you cannot fail is not a rating; it is reassurance.
3. What a prompt structurally cannot do
- It does not measure. QOVES says it uses "state-of-the-art landmarking technologies to map 521 facial points of your face" and assesses "more than 160 beauty tests" (qoves.com, read 2026-09-17). A general chatbot looking at a JPEG returns prose. If it states a canthal tilt in degrees or a gonial angle, no landmarking step produced that number.
- It is not reviewed. QOVES: "Every analysis is carefully reviewed by our team to ensure quality, consistency, and clarity." A prompt has no second reader.
- It may decline, by policy. Anthropic's usage policy prohibits using its models to "engage in behaviors that promote unhealthy or unattainable body image or beauty standards, such as using the model to critique anyone's body shape or size" (anthropic.com/legal/aup, read 2026-09-17). Whether a given prompt gets an answer, a hedge or a refusal is the provider's call and can change without notice.
- It keeps no baseline. A report or a scan you can repeat next month is a measurement series. A chat is a conversation.
4. Where SoftMaxx sits (ours, so read it as such)
SoftMaxx is AI, not a human analyst, and it is not a substitute for what QOVES sells. It differs from a bare prompt in two ways that the QOVES study makes relevant. First, geometry is measured rather than described: 478 landmark points are computed in your browser and sent to our backend as numeric ratios, as the methodology page sets out. Second, scoring is calibrated "so that population average sits near 5.0/10" — and you can check that against live data: across the 303 scans on record on 17 September 2026 the mean score was 5.0 and the median 4.8 (read from our public /api/stats endpoint; dated snapshots of the full distribution are published here). Low scores are used. That is the property the chatbots in the QOVES study lacked.
The limits are just as real: a vision model still reads the photo for everything geometry cannot capture, no person reviews your result, and the same methodology page says plainly that it "does not replace clinical evaluation."
5. Which to use for what
- Choosing between two photos, or a quick gut check: a chatbot prompt is fine. The ranking signal is real and it costs nothing.
- A number you intend to track, or sub-scores by feature: use a tool that measures and is calibrated to use the whole scale. The SoftMaxx front scan is free; see also face analysis apps compared.
- A written, human-reviewed protocol, and you can wait: that is QOVES's product — see what it costs and whether it is worth it.
- Anything surgical or medical: none of the three. A qualified clinician examines you in person.
6. What this page cannot tell you
The study is QOVES's own, published on its site; we found no journal version, and a vendor has an interest in showing that a free chatbot is not a replacement for its paid service. The face set is 102 neutral studio portraits, which is not how anyone photographs themselves for a prompt. We did not run prompts ourselves, so nothing here describes how a specific model will respond to your photo today. And attractiveness ratings, human or machine, describe agreement among raters — not worth, and not a diagnosis. If rating yourself is causing distress, that is a reason to stop and talk to someone, not a reason to find a more accurate rater.
7. Common questions
Can ChatGPT do a QOVES-style facial analysis?
It can describe a face and rank photos roughly the way people do, but it does not landmark or measure anything and it is not reviewed by a person. QOVES's own study found ChatGPT's average rating was 4.41 out of 7 against a human average of 3.02 on the same 102 faces, with no model ever giving the lowest score.
Is there a QOVES prompt for ChatGPT or Claude?
People search for one, but there is no official QOVES prompt, and a prompt cannot reproduce the product: QOVES says it maps 521 facial points, runs more than 160 tests and has every analysis reviewed by its team. A prompt returns an unmeasured opinion.
How accurate is AI at rating attractiveness?
In QOVES's study the four models agreed with human rank order at Spearman ρ .58 to .76, which is meaningful, but their absolute ratings were inflated: means of 4.41 to 5.21 on a 1–7 scale versus 3.02 for 2,513 human raters.
Why do chatbots rate everyone above average?
The QOVES page reports the pattern — "they rate almost everyone above average and never give the lowest scores" — without proving a cause. Provider policies against critiquing people's appearance are one plausible contributor; Anthropic's usage policy, for example, prohibits using its models to critique anyone's body shape or size.
Will Claude or ChatGPT refuse to rate my face?
Sometimes. It depends on the provider's policy and the wording, and it can change without notice. Anthropic's published usage policy prohibits promoting unhealthy or unattainable beauty standards.
Is SoftMaxx just ChatGPT with a prompt?
No. Facial geometry is measured from 478 landmarks computed in your browser, and scoring is calibrated so the population average sits near 5 out of 10 — the live mean across 303 scans on 17 September 2026 was 5.0. It is still AI, with no human reviewer, and it does not replace clinical evaluation.
What is the cheapest way to get a real baseline?
A free measured scan gives you an overall score and category sub-scores you can repeat later. A chatbot is free too, but its score is inflated and not repeatable as a series. QOVES costs $150 a year and returns a human-reviewed report within 28 days.