Observed Signal · May 20, 2026 · Technical Release · Source: DEV Community · Impact: 2/5 · Sentiment: Neutral

Benchmarking LLMs as Weekly Stock Pickers

Executive Signal Summary

A developer launched 1rok, an open benchmark that runs seven frontier LLMs as automated weekly stock‑picking agents. Each model receives $100,000 in paper capital and identical tools, prompts, and data; every Monday the models generate portfolios and Alpaca executes trades in paper accounts. The pipeline uses a multi-agent structure (Macro, Screener, six Analysts, Orchestrator, Constructor) with separate run and execute commands to keep decisioning auditable. The live experiment began January 20, 2026; a public leaderboard and agent traces are available. The aim is to evaluate downstream decision‑making under uncertainty rather than optimize for alpha, and to surface comparative behavior (risk sizing, panic in drawdowns, concentration mistakes) across models including GPT-5.5, Gemini 3.1 Pro Preview, Grok 4.3, DeepSeek V4 Pro, GLM-5.1, Kimi K2.6 and MiniMax M2.7.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Public benchmark demonstrating multi-agent LLM decision-making on a real downstream task (weekly stock picking); useful for model comparison but not an industry‑shifting platform or major vendor policy change.

SIGNAL RADAR

Track DeepSeek Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Project name: 1rok with live leaderboard at investingbench.vercel.app and repo at github.com/achaljhawar/1rok.
  • Benchmark runs seven LLMs, each assigned $100,000 of paper capital and identical tools/data; the experiment started on 2026-01-20.
  • Pipeline executes weekly (every Monday at 09:45 ET) using a multi-agent flow: Macro, Screener, six Analysts, Orchestrator, and Constructor.
  • Broker interactions use Alpaca paper accounts; tooling returns typed JSON from handlers connected to Alpaca, Yahoo Finance, FRED, and Tavily.
  • Models under test: GPT-5.5 (OpenAI), Gemini 3.1 Pro Preview (Google), Grok 4.3 (xAI), DeepSeek V4 Pro, GLM-5.1 (Zhipu), Kimi K2.6 (Moonshot), MiniMax M2.7.
Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: DEV Community•Published: May 20, 2026
Original Coverage Title: “Which LLM is the best stock picker? I built a benchmark to find out.”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIApr 11, 2026

RealDataAgentBench: Benchmark Reveals LLM Agents' Statistical Blind Spots

A developer published RealDataAgentBench, an open-source benchmark that evaluates LLM agents on correctness, code quality, efficiency, and statistical validity using reproducible seeded datasets and automated scoring. The benchmark includes 23 tasks across EDA, feature engineering, modeling, statistical inference, and ML engineering, and the author reports 163+ experiments across ~10 models (examples: GPT-4o, Claude Sonnet, Grok, Gemini 2.5, Llama via Groq). Key findings: GPT-4o and Claude Sonnet score similarly overall, GPT-4o is significantly cheaper per task, Groq/Llama runs are fast and low-cost but sometimes lack statistical rigor, and the biggest failure modes are statistical validity and code quality. The project is hosted on GitHub with a live leaderboard and supports budget flags and Groq free-first-test support.

Read assessment
Large Language Models (LLM) & AIJun 8, 2026

Top Open-Source Coding LLMs — June 2026 Leaderboard

A June 8, 2026 roundup surveys the rapidly changing open-weight coding LLM landscape, highlighting several new or updated models and practical deployment guidance. Key entrants include MiniMax M3 (released June 1, 2026; vendor-reported top SWE-bench Pro score, weights pending), Z.AI's GLM-5.1 (April 2026; 754B MoE, MIT license, designed for long-horizon autonomous execution), Moonshot AI's Kimi K2.6 (1T params with reasoning-state preservation for local agentic workflows), Alibaba's Qwen3.6-35B-A3B (April 16, 2026; single-GPU local deployment, high SWE-bench Verified), DeepSeek V4 (April 24, 2026; V4-Flash self-hostable variant), and Codestral 22B (leader for IDE autocomplete with 95.3% FIM pass@1). The article emphasizes benchmark contamination (HumanEval saturation), recommends benchmark types that better discriminate agentic and long-horizon coding ability, and provides hardware and practical stacks for different developer and organizational needs.

Read assessment
Large Language Models (LLM) & AIJul 14, 2026

Author Tests 300+ LLMs and Ends Benchmark

A developer ran a long-running benchmark of large language models (LLMs) across real-world agent coding tasks, testing over 300 models reachable via OpenRouter and local hardware. The benchmark covered ten practical tasks, used a 400-token cap, low temperature, pattern-matching scoring, and pre-flight verification. The author retired the public leaderboard (published at 168 models) because rapid model churn, harness limitations, and lack of audience made continuous maintenance unjustified. The author retains the habit of ad-hoc testing, preserved the archived data for on-demand queries, and argues that static leaderboards quickly become stale as models improve.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.