Observed Signal · Jul 17, 2026 · Research Experiment · Source: t3n · Impact: 2/5 · Sentiment: Neutral
Researcher has LLMs simulate harbor cities over 6,000 years
Ethan Mollick, an AI researcher at the Wharton School (University of Pennsylvania), created an experimental benchmark called the AI Harbor Town Gallery that asks language models to generate code and simulate the growth of harbor cities from 3000 BCE to 3000 CE. The interactive site lets users play, pause and scrub simulations and shows Mollick's scoring for code quality and design alongside notes about attempts and errors. Tested models span older and newer LLMs (examples in the article include GPT-3.5 Turbo, GPT-4, various Claude variants and more recent GPT-5.6 Sol Pro). According to the write-up, Claude 4.8 Opus Extra performed best in these tasks, while GPT-3.5 Turbo and GPT-4 performed worst in this specific benchmark.
Shows comparative LLM capabilities in extended code generation and simulation tasks — relevant to generative-AI use cases (creative/code) but not a major platform policy or industry-shifting announcement for AdTech/MarTech.
Track t3n Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Ethan Mollick of the Wharton School created the AI Harbor Town Gallery benchmark to have language models build and simulate harbor cities over a period from 3000 BCE to 3000 CE.
- The benchmark evaluates models on generated code quality and overall design, and reports how many attempts and errors occurred for each output.
- Tested models mentioned include GPT-3.5 Turbo, GPT-4, Claude variants (including Claude 4.8 Opus Extra and Claude Fable 5) and GPT-5.6 Sol Pro.
- Claude 4.8 Opus Extra achieved the best results in the harbor-city simulations; GPT-3.5 Turbo and GPT-4 performed poorest in this test.
- Simulations are available on the project website and can be controlled via an interactive timeline (start, stop, fast-forward).
Connected Companies & Entities
7 Entities mapped“The article appears on t3n.de (t3n – digital pioneers)....”
“The article notes: 'Here you can find external content from X Corp. that complements our editorial offering on t3n.de.'...”
“The article includes the editorial note: 'Here you can find external content from Podigee GmbH that complements our editorial offering on t3...”
“The article includes the editorial note: 'Here you can find external content from TargetVideo GmbH that complements our editorial offering o...”
“The tested models range back to GPT-3.5 Turbo, which was released in March 2023 (models such as GPT-3.5 Turbo and GPT-4 are referenced in th...”
“The article references Claude model variants, noting that Claude 4.8 Opus Extra performed best and that Claude Fable 5 was also tested....”
“The piece states that models including GPT, Claude and Gemini are asked to simulate harbor cities over several millennia....”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
LLM Leaderboard: Top AI Models (April 2026)
A benchmarking roundup (published Apr 23, 2026) ranks the leading large language models across multiple independent systems. LM Arena’s human-preference Elo list places Claude Opus variants at the top, with claude-opus-4-7 (1504 Elo) leading. Claude Opus 4.7 also tops coding benchmarks (82.0% on SWE-bench Verified). The Artificial Analysis Intelligence Index shows a three-way tie (score 57) between Claude Opus 4.7, Google’s Gemini 3.1 Pro Preview, and OpenAI’s GPT-5.4. The report highlights price-performance tradeoffs: DeepSeek V3.2 offers the lowest input cost ($0.29 per million tokens), while Kimi K2.6 (Moonshot AI) is the highest-profile open-weight model with a 256K context window. The article explains ranking methodologies (LM Arena, SWE-bench Verified, GPQA Diamond, composite index) and gives model recommendations by use case (coding, long context, high-volume, self-hosted).
Author Tests 300+ LLMs and Ends Benchmark
A developer ran a long-running benchmark of large language models (LLMs) across real-world agent coding tasks, testing over 300 models reachable via OpenRouter and local hardware. The benchmark covered ten practical tasks, used a 400-token cap, low temperature, pattern-matching scoring, and pre-flight verification. The author retired the public leaderboard (published at 168 models) because rapid model churn, harness limitations, and lack of audience made continuous maintenance unjustified. The author retains the habit of ad-hoc testing, preserved the archived data for on-demand queries, and argues that static leaderboards quickly become stale as models improve.
Open Models Closing Capability Gap with Frontier AI
SemiAnalysis presents an analysis showing open-source AI models have closed the capability gap with closed-source frontier models faster with each successive era of LLM development. Using curated benchmarks across three eras (early scaling, reasoning, agentic) and evaluation tooling (Prime Intellect), the author finds a consistent pattern: open models take roughly half as long each generation to match the first closed-source model of that era. The piece cites specific model milestones (e.g., Llama releases, DeepSeek R1, o1-preview, GLM and Kimi variants), usage statistics (Fireworks processing ~40T tokens/day), and commercial impact (Anthropic’s Claude Code contributing to >$65B ARR). The article highlights benchmark limitations and productization (model + harness) as important factors beyond raw benchmark scores.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
