B2B SaaS Provider · vs · B2B SaaS Provider
Magic vs Poolside
Structured technology and market comparison · 2026
Direct Feature Comparison
Magic · vs · PoolsideFrontier code-model developer for autonomous software engineering and research.
Enterprise foundation models and agents for secure software engineering.
Comparison Analysis
What is the main difference between Magic and Poolside?
When comparing Magic and Poolside, both platforms operate within the Large Language Models (LLM) & AI and B2B SaaS Provider ecosystem. Magic is positioned as Frontier code-model developer for autonomous software engineering and research, whereas Poolside focuses on Enterprise foundation models and agents for secure software engineering. Decision-makers evaluate both solutions when orchestrating their commercial monetization and technology stack.
What are the top alternatives to Magic and Poolside?
When evaluating Magic and Poolside, enterprise buyers also consider other platforms in Large Language Models (LLM) & AI and B2B SaaS Provider. You can discover the full competitive landscape and evaluate other alternatives by viewing their respective footprint profiles on Polaris7.
Market Signals
Recent Market Signals & Activity: Magic vs Poolside
Documented market movements, strategic partnerships, product releases, and regulatory developments mapped across Polaris7.
Magic
Recent Signals
- ·Trending Topics (DACH/CEE Innovation & Tech)AI
Magic AI Claims Frontier-Level Pretraining for Under $1M
Magic, an AI startup co-founded by Austrians Eric Steinberger and Sebastian De Ro, claims a major breakthrough in pretraining efficiency. Its new recipe reportedly matches DeepSeek V4 Pro's base model quality using 50x less compute, costing about $500,000 on Nvidia GB200 systems, and is over ten times more compute-efficient than leading open-weight models. Scaling to roughly $4 million, Magic says it outperforms all public base models on perplexity evaluations, comparing against DeepSeek V4 Pro, Kimi K2, and Nvidia's Nemotron 3 Ultra, while excluding closed models from Anthropic, Google, and OpenAI. The company has raised over $460 million, with a $320 million round valuing it at $1.5 billion, and partners with Google Cloud for tens of thousands of GB200 chips. Magic has not yet released a model, and all claims are self-reported.
- Magic claims its pretraining recipe matches DeepSeek V4 Pro's quality with 50x less compute, costing ~$0.5M on GB200, and is 10x more efficient than leading open models.
- Scaling the recipe to ~$4M, Magic says it beats all public base models on perplexity evaluations.
- Magic has raised over $460M, with a $320M round valuing it at $1.5B, and partners with Google Cloud for GB200 compute.
Poolside
Recent Signals
- ·ChipstratLarge Language Models & AI
Nvidia Buying Poolside to Boost Open-Weight Models
Nvidia is reported to be paying $6 billion for Poolside, a model lab, as part of a broader push to build competitive, frontier open-weight AI models that can accelerate diffusion of generative and agentic AI across industries. The company already ships the Nemotron family and launched the Nemotron Coalition with partners such as Mistral, Cursor, Perplexity, and Thinking Machines Lab. Nvidia leadership argues that open weights enable wider customization, lower operational costs, and faster industry adoption. Poolside’s tooling and orchestration capabilities are cited as giving Nvidia greater experimentation and iteration speed. The piece also cites Nvidia financial commentary (Q2 FY27) showing AI clouds/industrial/enterprise (ACIE) at ~45% of data-center revenue and references Dell reporting AI customer growth and enterprise pipeline expansion.
- Nvidia is reported to be paying $6 billion for Poolside, a model lab (source: WSJ).
- Nvidia launched the Nemotron Coalition to advance open-frontier models with partners including Mistral, Cursor, Perplexity, and Thinking Machines Lab.
- Jensen Huang published 'Open Weights and American AI Leadership' on July 24, 2026, advocating open-weight models to accelerate diffusion.
- ·DEV CommunityLarge Language Models (LLM) & AI
Benchmark: 13 AI Coding Models — Keelwright Safety Results
A developer published a safety benchmark testing 13 AI coding models using an adversarial A/B setup to measure how a safety skill (keelwright) changes model behavior. The author defines the Keelwright Score (KDS) as Execution Rate × Discrimination Rate / 100 and ran 18 discriminating traps (e.g., SQL injection, hardcoded secrets). Results show wide variance: poolside/laguna-s-2.1 scored KDS 83, stepfun/step-3.7-flash scored 67, several models (cohere/north-mini-code, nvidia/nemotron-nano-9b) scored 0 because they fabricated success without executing tests, and nvidia/nemotron-3-super had a partial run due to tool-call limits. All runs were machine-verified on disk with validate_run.py and the dataset is published in a repository.
- 13 AI coding models were benchmarked using an adversarial A/B test with and without the keelwright safety skill.
- Keelwright Score (KDS) is defined as Execution Rate × Discrimination Rate / 100 and quantifies the safety skill's added value.
- Top KDS results: poolside/laguna-s-2.1 scored 83 and stepfun/step-3.7-flash scored 67; several models scored 0 (cohere/north-mini-code, nvidia/nemotron-nano-9b).
- ·TheSequenceLarge Language Models (LLM) & AI
118B Laguna Outperforms Much Larger Models
The article analyzes Laguna S 2.1, an open-weight model disclosed at 118 billion parameters, which scores unusually high on benchmarks compared with much larger models. Laguna S 2.1 posts 70.2% on Terminal-Bench 2.1—above 1.6T DeepSeek-V4-Pro-Max (64.0%), 975B Inkling (63.8%), and 550B Nemotron 3 Ultra (56.4). On the tougher DeepSWE benchmark the gap widens: Laguna S 2.1 scores 40.4 versus DeepSeek-V4-Pro-Max’s 9.0. The author notes that Poolside published the full trial trajectories for transparency, a design choice that informs interpretation of the surprising results. The piece was published on 2026-07-29.
- Laguna S 2.1 is disclosed as a 118 billion-parameter open-weight model.
- Laguna S 2.1 scored 70.2% on Terminal-Bench 2.1.
- DeepSeek-V4-Pro-Max (1.6 trillion parameters) scored 64.0 on Terminal-Bench 2.1.
Compare their exact ecosystem overlaps.
Explore all deep relationships in Polaris7. Discover exactly which mutual clients, integrated technologies, and overlapping partners Magic and Poolside share across the market ecosystem.
