Observed Signal · Jul 21, 2026 · Case Study / Experiment · Source: UX Collective · Impact: 2/5 · Sentiment: Positive

AI speeds UX review but doesn't replace expert

Executive Signal Summary

A UX practitioner tested two paid LLMs — Claude Fable 5 and ChatGPT-5.6 Sol — by feeding them a full UX brief, research, Google Analytics exports and Clarity heatmaps to produce an expert UX review. Both models produced useful observations and some false or irrelevant findings; Claude produced a single-document output and was subjectively more accurate. The author estimated the workflow was roughly 15% faster with AI, and ultimately assembled the final analysis with Claude, using the model mainly for drafting, grammar, and consistency. The piece argues AI is a strong first-pass partner when grounded in real data and expert heuristics (citing Baymard research and a 2025 GPT-4o study), but current LLMs do not yet replace human judgment for severity, prioritization, and context-sensitive evaluation. The author plans to build a reusable agent skill encoding professional heuristics.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Practitioner case study showing practical LLM usage for UX reviews—useful for product/UX teams and MarTech workflows but not industry-shifting. Reinforces best practice of grounding models in expert heuristics.

SIGNAL RADAR

Track Google Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • The author tested two paid LLMs in parallel: Claude Fable 5 and ChatGPT-5.6 Sol.
  • Inputs provided to the models included the brief, industry research, Google Analytics exports, and Clarity heatmaps/attention & scroll maps.
  • The author estimated the AI-assisted workflow was roughly 15% faster than doing the work entirely by hand.
  • Both models produced useful observations plus irrelevant or non-existent problems; Claude was subjectively more accurate and could produce the result in a single document.
  • The article cites a 2025 study reporting GPT-4o found roughly 21% of the usability problems identified by human experts while also raising some false problems.

Connected Companies & Entities

7 Entities mapped

“I had a set of questions about the target audience, the goals, and the USP, an industry research document, plus access to Google Analytics a...”

“I had a set of questions about the target audience, the goals, and the USP, an industry research document, plus access to Google Analytics a...”

“I worked with two models in parallel: Claude Fable 5 and ChatGPT-5.6 Sol, both on paid subscriptions, of course....”

“I worked with two models in parallel: Claude Fable 5 and ChatGPT-5.6 Sol, both on paid subscriptions, of course....”

“According to a [late-2024 Reuters piece](https://www.reuters.com/technology/artificial-intelligence/openai-rivals-seek-new-path-smarter-ai-c...”

“According to a [late-2024 Reuters piece](https://www.reuters.com/technology/artificial-intelligence/openai-rivals-seek-new-path-smarter-ai-c...”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: UX Collective•Published: Jul 21, 2026
Original Coverage Title: “I handed a UX review over to AI. Here’s what happened.”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Web/App Development & UX DesignJun 15, 2026

T-shaped UX Becomes AI‑Empowered Polymath Architect

A Medium analysis by Patrick Neeman (published 2026-06-15) argues that generative AI is collapsing the handoffs that made the traditional T-shaped UX model valuable, accelerating a return to polymathic practitioners who can direct tools across the full product stack. The piece cites industry surveys showing widespread AI adoption among designers and developers, reports UX staffing reductions, and recommends concrete actions: map handoffs, ship work outside your lane using AI, encode judgment as prompts/checklists, and protect deep specialization where quality still matters. The author positions AI as an "assembly line" that automates production and handoffs while elevating human judgment, context, and taste as the enduring sources of value.

Read assessment
Large Language Models (LLM) & AIApr 29, 2026

Thoughtful AI for UX Research Leaders

Ashlee Edwards, Ph.D. outlines a measured approach to introducing LLM- and neural‑network‑based AI tools into UX research teams. Writing from the perspective of a research leader managing an eight-person team, she recommends setting a clear "north star" that AI should support—not replace—research quality, defining skill-preserving guidelines (e.g., avoid using AI to craft research questions; allow AI for data cleaning with human review; label AI-generated content), and framing tool adoption with risk-vs-reward questions about output quality, time savings, and cost-effectiveness. She describes implementing a NotebookLM instance for searchable research sources, pressure-testing prompts, tracking usage and outcomes, and documenting benefits and tradeoffs to ensure quality and ROI. The piece advocates transparency, documentation, and continued human oversight when deploying AI in research workflows.

Read assessment
Large Language Models (LLM) & AIJul 17, 2026

AI Feedback Loops Make Agents More Useful

The author argues that the newest generation of large language models (LLMs) and surrounding tooling have reached a practical threshold: they can operate browsers, connect to many data sources, and automate recurring analysis. The author describes a working example: an agent (built on Fable) that weekly scans academic papers and improves selection by ingesting the author’s audible, in-the-moment reactions captured with Wispr Flow. Giving the agent these revealed-preference signals allowed it to self-adjust and produce materially better results than earlier models (Opus 4.8, GPT-5.5). The piece recommends building simple agentic feedback loops (using ChatGPT or Claude) for recurring reports and warns that connecting models to private data is what makes them particularly powerful — and “spooky.”

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.