Observed Signal · Oct 14, 2024 · Research Publication · Source: Trending Topics · Impact: 3/5 · Sentiment: Negative

Apple Study: LLMs Can't Calculate or Reason Properly

Executive Signal Summary

Apple researchers published a paper examining the limits of mathematical reasoning in large language models and introduced a new benchmark, GSM-Symbolic. The team tested state-of-the-art models, including OpenAI's and Meta's Llama 3-8b, by adding contextual details to math problems. Accuracy fell by up to 65 percent when a seemingly relevant sentence was inserted, and performance also dropped when only numeric values changed. Apple argues that the existing GSM8K benchmark, co-developed by OpenAI with Surge AI, is insufficient for measuring true reasoning. The researchers found no evidence of formal reasoning in LLMs and concluded that reliable AI agents cannot be built on models that merely reproduce patterns.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Apple, a major platform, provides evidence that LLMs lack formal reasoning, which raises concerns about the reliability of AI agents and automated decision-making being adopted across AdTech and MarTech.

SIGNAL RADAR

Track Apple Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Apple researchers published the paper 'Understanding the limitations of mathematical reasoning in large language models' and introduced the GSM-Symbolic benchmark.
  • Accuracy of the tested LLMs declined by up to 65 percent when irrelevant but plausible context was added to math problems.
  • Apple found no evidence of formal reasoning in language models, including OpenAI's models and Meta's Llama 3-8b.
  • GSM8K, the basis for GSM-Symbolic, contains 8,000 grade-school math problems and was co-developed by OpenAI and Surge AI.
  • Apple concluded that reliable AI agents cannot be built while LLMs only reproduce patterns.

Connected Companies & Entities

3 Entities mapped

“A group of AI researchers at Apple published the paper 'Understanding the limitations of mathematical reasoning in large language models'....”

“The mention of some smaller kiwis led Apple to note that, among others, OpenAI's models and Meta's Llama3-8b drew wrong conclusions....”

“The mention of some smaller kiwis led Apple to note that, among others, OpenAI's models and Meta's Llama3-8b drew wrong conclusions....”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: Trending Topics•Published: Oct 14, 2024
Original Coverage Title: “Apple-Studie: LLM-basierte AI-Modelle können nicht richtig rechnen und denken”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

AIOct 9, 2026

Unlocking Claude's Full Value: Plans, Workflows, and Setup

This article analyzes the economics of Anthropic's Claude subscription plans, comparing them to OpenAI's offerings. SemiAnalysis found a $200 Claude Max plan provides about $11,700 worth of Claude Opus 5.5 usage at API prices, versus roughly $2,100 for OpenAI's equivalent. Most subscribers use only a small fraction of their allowance, making the plans profitable for Anthropic. Recent product changes, including the merger of Cowork into Claude chat, new model versions (Opus 5.5, Sonnet 5.5, Haiku 5.5), and monthly API credits on Max and Team plans, make it easier to use the full allowance. The article provides a comprehensive guide with model routing tables, a map of the Claude stack, a 5-layer operating system, and 17 workflows to help users maximize their subscription value.

Read assessment
Social PlatformOct 9, 2026

Threads tests Gems to honor valuable posts daily

Threads is testing a new feature called Gems, which awards up to 25 outstanding conversations daily with a visible diamond badge. The selection is algorithm-based, focusing on originality, timeliness, and authentic discussion, regardless of follower count. Posts or direct replies qualify, but reposts and content cross-posted from Instagram are excluded. The feature is currently limited to public US-based accounts of users aged 18 and older, though the badges are visible globally, including in Germany. Recipients get a notification and can share the badge. Users can hide the label if desired. Initial hints appeared in August via app researcher Alessandro Paluzzi. Threads aims to encourage original and engaging content, with no negative impact on reach for non-recipients.

Read assessment
Social Media AdvertisingOct 9, 2026

Meta Bans TikTok Ads on Its Platforms in Multiple Regions

Meta has banned advertisements and paid marketing messages for TikTok, owned by ByteDance, on its platforms in the US, Canada, Egypt, Indonesia, Japan, Thailand, and Vietnam, effective October 8, 2026. The ban applies to both direct ByteDance ads and third-party campaigns linking to TikTok or other ByteDance services. Meta spokesperson Chris Sgro confirmed the move, citing standard practice to avoid supporting competitors. This intensifies rivalry between the social media giants, both designated as gatekeepers under the EU's Digital Markets Act. The ban excludes EU countries, but raises questions about cross-platform advertising. It follows Meta's $16.7 billion settlement with US states over youth safety, which requires TikTok and YouTube to implement similar safeguards, suggesting a strategic pressure tactic.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.