Claude beats GPT—but isn't convincing yet

A study by Wortliga, commissioned by Sistrix, compared 2,112 B2B marketing texts generated by Claude, Gemini, and GPT. The result is clear: Without precise prompts, none of the models writes clearly enough—and that poses risks for companies.

Sentence length affects the clarity of marketing texts. Source: zvg

AI language models are considered efficient tools for marketing and sales. But how good are they, really? A study by Wortliga, commissioned by the SEO analytics provider Sistrix, provides the first systematic comparison of the flagship models from Anthropic, Google, and OpenAI—and reaches a sobering conclusion.

2,112 texts, eleven genres, three models

The analysis included 2,112 B2B texts from eleven text genres, including social media posts, promotional emails, SEO blog articles, case studies, and white papers. The quality of the texts was measured using the WORTLIGA score, which evaluates readability, linguistic barriers such as nested sentences and passive voice, as well as tone and substance. Texts are considered understandable only if they score 60 points or higher.

GPT 5.5 ranks last in the WORTLIGA rankings. Source: zvg

The result: Anthropic’s Claude Opus 4.7 took first place and delivered the most consistent language performance across all industries—with an average score of 47.7 points. Google’s Gemini 3.1 Pro (preview version) followed in second place with 46.8 points, but stood out for a noticeable abundance of clichés and filler words. OpenAI’s GPT 5.5 model came in last with just 37.7 points. Without precise guidelines, it tends to use bureaucratic language and complex technical jargon.

The uncontrolled use of AI is risky

Study director Gidon Wagner of Wortliga draws a clear conclusion: «The results show that uncontrolled AI-driven communication in marketing and sales is risky.» Not a single one of the three models reached the 60-point threshold for comprehensible text—even the best model fell far short of that mark.

No AI model falls within the green comprehensibility range. Source: zvg

For Johannes Beus, founder and CEO of Sistrix, this goes beyond mere questions of style: «Whether a text is interpreted correctly by an AI system depends, among other things, on its linguistic clarity and structure. Language is no longer a matter of taste. So what was long considered a »nice-to-have’ is now a key quality criterion, both from an accessibility standpoint and in terms of machine-readable content.”

The Chameleon Effect in Prompting

Another key finding of the study is the so-called “chameleon effect”: If a prompt uses a distant, formal style, the models inevitably adopt it—and comprehensibility drops dramatically. On average, the texts scored only 4.4 points under such conditions. Marketing emails then read like official correspondence. This behavior was particularly pronounced in GPT 5.5.

Claude is the best example; GPT is a total failure. Source: zvg

When given vague prompts, the systems even ignored the rules for certain types of text. This shows that it is not the model alone that determines quality—but rather the quality of the instructions it is given.

People Remain Indispensable

The study notes that AI models cannot, on their own, write compelling B2B copy. Gidon Wagner sums it up: «The data shows that AI models are inherently incapable of writing compelling B2B copy.» Johannes Beus adds in the study’s foreword: «It is an honest assessment of what artificial intelligence is actually capable of today and where humans remain indispensable.»

AI-generated texts are highly readable but lack depth. Source: zvg

 

More information: https://wortliga.de/

(Visited 182 times, 1 visits today)

More articles on the topic