Observed Signal · Jul 24, 2026 · Technical Release · Source: Meedia · Impact: 2/5 · Sentiment: Neutral
Havas finds LLMs fall short in media planning
Havas Media Germany published the Havas AI Media Quality Index (HAI-Q), a benchmark that tests how well AI systems perform concrete media-planning tasks. Using a 35-question testset tailored to the German market, Havas reports that general large language models often produce convincing-sounding but unreliable answers, particularly on data-driven decisions and numerical calculations. Havas’ internally developed “Agentic Media Machine” scored 33/35 (94.3%), while third-party LLMs scored substantially lower (GPT-5: 16/35; Claude Sonnet 4.5: 6/35; GPT-4o: 4/35). Havas frames the results as evidence for the need for specialised systems, proprietary data and domain expertise rather than a condemnation of AI, and says it will continue running and expanding the benchmark internationally.
Agency-published benchmark highlights limitations of general LLMs for concrete media-planning tasks and signals industry interest in specialised AI systems; relevant to agencies, advertisers and MarTech but not an industry-shifting platform policy or major platform technical change.
Track MEEDIA Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Havas Media Germany created the Havas AI Media Quality Index (HAI-Q) as a benchmark for AI in media planning.
- HAI-Q used a testset of 35 tasks tailored to the German market to evaluate media-planning performance.
- Havas’ Agentic Media Machine scored 33 out of 35 (94.3%) on the HAI-Q.
- GPT-5 scored 16/35 (45.7%); Claude Sonnet 4.5 scored 6/35 (17.1%); GPT-4o scored 4/35 (11.4%).
- Havas found general LLMs answer typical media-planning questions correctly only about half the time and showed notable weaknesses on data-based decisions and numerical calculations.
Connected Companies & Entities
10 Entities mapped“The article was published on MEEDIA and references prior MEEDIA reporting about AI in media planning....”
“Google appears in the cookie/statistics section as the creator of analytics cookies (_ga) used on the site....”
“LinkedIn appears in the cookie/marketing table (e.g., cookie 'bcookie' listed with LinkedIn as creator)....”
“DataReporter GmbH is listed in the cookie table as the creator of consent-related cookies (e.g., _webcare_consentid)....”
“CloudFlare is listed in the cookie table as the creator of the '__cf_bm' cookie....”
“WordPress is listed in the cookie table (e.g., 'wordpress_test_cookie') as a creator of session cookies for site functionality....”
“YouTube is listed as the creator of an embedded-content consent cookie ('susc_social_embed_consent')....”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
AI Can't Spot AI Ads but Can Judge Human Creative Quality
Andrew Tindall, a senior leader at System1, explores whether AI can effectively evaluate human-made advertising, given that humans cannot reliably identify AI-generated ads. He cites four major studies over ten months, including NYU Stern's 'AI Advertising Paradox' (AI ads outperformed human ads by up to 19% without disclosure, but performance dropped 31.5% when disclosed) and a Taboola study with 500M+ impressions finding similar results for short-term metrics like CTR. In his own experiment, he ran ten well-known human-made ads through System1's AI testing tool, comparing its predictions to human-based ratings. Results show the AI predicted the Fluency Rating correctly 10 out of 10 times and the Star Rating 9 out of 10 times, with one refusal to score. AI performed best at extreme high and low performers, struggled with mid-table ads, and highlighted the importance of acknowledging uncertainty. Tindall concludes that AI is a probabilistic prediction machine, better used to supplement rather than replace human judgment in creative testing.
Marketers scramble to measure AI brand visibility
Marketers are increasingly concerned about how AI chatbots and large language models (LLMs) cite or mention their brands in responses, prompting a shift in advertising strategies. The Interactive Advertising Bureau (IAB) is developing a framework to standardize the measurement of AI visibility. Data shows brand citation shares vary significantly across platforms like Microsoft Copilot, ChatGPT, and Gemini, while LLMs show preferences for different source types. Trust in AI search is growing, with 95% of US respondents finding AI answers as trustworthy as search engines. However, measurement is challenging due to the non-deterministic nature of AI models. This has led to new roles like Head of AI Search and the allocation of ad budgets to influence AI recommendations.
AI in Media Buying: Expectations Outpace Experience
A new IAB Europe study reveals that while 58% of industry professionals expect agentic ad buying to become mainstream within a year, actual adoption lags significantly. Of 47 respondents, 36 have either not yet deployed agentic systems or use technologies where humans still control planning and execution. The 'Impact of AI on Digital Advertising Report 2026', based on 50 interviews, shows that AI is widely used for reporting and campaign analysis (86% overall), but only 11 out of 29 detailed respondents use it for agentic buying and selling. A complementary survey by StackAdapt (500 marketers, 6 countries) found similar patterns, with half citing fragmented data pipelines as a barrier. Companies are cautious about granting full autonomy, with most preferring human-in-the-loop models. The average performance rating for AI in operational advertising was only 2.76 out of 5. The industry is still defining governance, training, and integration practices to bridge the gap between expectations and real-world experience.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
