Observed Signal · Jul 7, 2026 · Study · Source: Meedia · Impact: 3/5 · Sentiment: Negative

Study: Claude Tops Gemini in Marketing Texts

Executive Signal Summary

A Wortliga study commissioned by Sistrix compared current large language models from Anthropic, Google and OpenAI on the quality of marketing and sales copy. Using 2,112 B2B texts across 11 text types (e.g., social posts, sales emails, SEO articles), researchers evaluated readability, sentence structure, passive voice and style. Claude Opus 4.7 scored highest (47.7 Wortliga points), Gemini 3.1 Pro (pre-release) scored 46.8, and GPT 5.5 scored 37.7; Wortliga defines texts as "understandable" only from 60 points. The study used the models' APIs, neutral temperature settings and eight different prompts; authors warn that models often produce less-understandable, overly formal text without precise prompts, indicating current LLM limitations for autonomous B2B marketing copy.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Comparative evaluation of leading LLMs for marketing copy highlights current limitations in automated B2B content quality — relevant to marketers, MarTech vendors and agencies planning to adopt LLMs for content creation.

SIGNAL RADAR

Track Anthropic Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Wortliga conducted a study for Sistrix comparing language models from Anthropic, Google and OpenAI.
  • The evaluation used 2,112 B2B marketing and sales texts across 11 text genres (e.g., social media posts, sales emails, SEO articles).
  • Claude Opus 4.7 achieved the highest average Wortliga score of 47.7; Gemini 3.1 Pro (pre-release) scored 46.8; GPT 5.5 scored 37.7.
  • Wortliga defines texts as "understandable" only from 60 points onward.
  • Methodology: models were accessed via their APIs with neutral temperature settings and eight different prompts to ensure comparable conditions.

Connected Companies & Entities

4 Entities mapped

“Eine Studie von Wortliga im Auftrag von Sistrix hat die aktuellen Sprachmodelle von Google, OpenAI und Anthropic miteinander verglichen....”

“Eine Studie von Wortliga im Auftrag von Sistrix hat die aktuellen Sprachmodelle von Google, OpenAI und Anthropic miteinander verglichen....”

“Eine Studie von Wortliga im Auftrag von Sistrix hat die aktuellen Sprachmodelle von Google, OpenAI und Anthropic miteinander verglichen....”

“Frank Puscher 7. Juli 2026 um 13:14 (article published on MEEDIA)....”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: Meedia•Published: Jul 7, 2026
Original Coverage Title: “Studie vergleicht Sprachmodelle für Marketing- und Vertriebstexte - Claude gewinnt knapp vor Gemini. ChatGPT mit Rückstand.”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIAug 26, 2026

Study: Google Gemini Leads for Academic Writing

A Studyarena analysis of 6,851 anonymous college-student votes (August 2026) compared AI-generated academic writing from Google's Gemini, Anthropic's Claude, and OpenAI's ChatGPT. In a blind test, Gemini ranked first with a 39.6% selection rate, followed by Claude at 31.8% and ChatGPT at 29.2%. The study found that students preferred longer answers (selected responses were on average 37% longer) and lower 'reasoning' levels often scored higher (40.7% vs. 29.5% for high-reasoning answers). Additionally, a Hochschule Darmstadt survey reported that 92% of German students use AI tools in their studies, up from 63% in 2023. Studyarena recommends using normal or low reasoning settings for writing tasks, while noting that different models excel at different tasks: ChatGPT for research, Claude for planning, and Gemini for writing. The leaderboard remains live for further voting.

Read assessment
Large Language Models (LLM) & AIApr 23, 2026

LLM Leaderboard: Top AI Models (April 2026)

A benchmarking roundup (published Apr 23, 2026) ranks the leading large language models across multiple independent systems. LM Arena’s human-preference Elo list places Claude Opus variants at the top, with claude-opus-4-7 (1504 Elo) leading. Claude Opus 4.7 also tops coding benchmarks (82.0% on SWE-bench Verified). The Artificial Analysis Intelligence Index shows a three-way tie (score 57) between Claude Opus 4.7, Google’s Gemini 3.1 Pro Preview, and OpenAI’s GPT-5.4. The report highlights price-performance tradeoffs: DeepSeek V3.2 offers the lowest input cost ($0.29 per million tokens), while Kimi K2.6 (Moonshot AI) is the highest-profile open-weight model with a 256K context window. The article explains ranking methodologies (LM Arena, SWE-bench Verified, GPQA Diamond, composite index) and gives model recommendations by use case (coding, long context, high-volume, self-hosted).

Read assessment
Large Language Models (LLM) & AIMar 7, 2026

GPT-5.4 Shows Strengths and Unexpected Failures

The author conducted six structured, blind evaluations comparing OpenAI’s GPT-5.4 (positioned for professional workflows) to Claude (Opus 4.6) and Google’s Gemini 3.1. Results show GPT-5.4 outperforms on certain professional tasks—quantitative modeling, file processing and self-knowledge—yet it produced confidently wrong answers on simple real-world questions where other frontier models succeeded. The analysis highlights a notable failure mode the author terms the “pipeline problem,” discusses a product-level split in model behavior, and interprets OpenAI’s direction as building agentic infrastructure rather than a traditional chatbot. The piece argues models are converging in raw capability but diverging in product philosophy, urging readers to focus on what benchmarks measure rather than just who wins them.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.