Baseten
Baseten is a b2B platform for AI model serving, inference and deployment.
Analyst Perspective
Baseten is a private B2B software company that provides infrastructure for deploying, serving, scaling and managing machine learning models in production. Its platform is built around production inference workloads, with offerings spanning managed cloud deployments, self-hosted and hybrid environments, model APIs, workflow orchestration, model management and training. The company sells primarily to developers, machine learning engineers, AI platform teams and enterprise engineering organisations that need low-latency, compliant and cost-efficient model operations. The business monetises through a mix of software subscription, enterprise contracts and usage-based infrastructure billing. Token-based pricing is used for managed model APIs, while deployment products are billed on underlying compute consumption such as GPU, CPU and memory, with premium enterprise pricing for dedicated environments, SLAs and compliance needs. Recent funding rounds and the acquisition of Parsed indicate an effort to broaden from inference serving into a more complete AI platform covering post-training and model lifecycle workflows.
Analyst Signal Briefing
Updated: 4 Aug 2026Baseten has finalised a $1.5 billion Series F funding round, reaching a $13 billion valuation to scale its inference infrastructure. Maintaining its strategic role as a primary provider for Microsoft’s MAI model family, the company is further specialising in "inference engineering." This involves productionising open-weight models like GLM-5.2 through quantisation and speculative decoding to optimise throughput and latency. These technical advancements, alongside ecosystem integrations via Azure Foundry and GitHub Copilot, reinforce Baseten's position as a critical layer for enterprise multi-model adoption and high-performance, cost-effective model serving.
Explorer Tier
Start exploring for free
Start with public company intelligence. Save companies, build your first watchlist, and unlock deeper strategic insights when you are ready.
- View public Company Profiles
- Save/watch companies
- Build your first Watchlist
- Access additional market signals
Key insights about Baseten
Category Differentiation
This company is not a foundational model developer or a general-purpose public cloud provider. It is an AI infrastructure and model serving platform focused on deploying and operating models for business customers.
Baseten: About
Baseten creates value by abstracting the operational complexity of production AI systems for business customers. It provides a proprietary platform that lets technical teams deploy models, access managed LLM endpoints, orchestrate inference pipelines, manage model lifecycle workflows and choose between shared cloud, dedicated, self-hosted or hybrid deployment modes. This reduces infrastructure burden, improves performance and compliance, and helps customers control GPU utilisation and inference cost. Revenue is generated from recurring platform access, enterprise infrastructure commitments and metered usage tied to tokens and compute consumption.
How Baseten Works & Monetises
Business model analysis and core revenue streams
Baseten uses a hybrid commercial model combining SaaS-style platform access with consumption-based billing. Managed model APIs are priced per million tokens, making this a pay-per-use API revenue stream. Cloud and deployment offerings monetise through compute-based usage charges across GPU, CPU and memory. Enterprise tiers add custom contracts, higher limits, dedicated infrastructure, SLAs, compliance features and negotiated volume discounts. A free or entry tier supports developer adoption, while larger committed workloads expand into higher-value enterprise agreements.
Revenue Channels
Products & Services in Categories
Verified structural categorizations from the graph
Technology
Recent Signals (Baseten)
Baseten Guests Discuss Inference Engineering Advancements
A long-format interview (published 2026-08-03) features Baseten's Philip Kiely and Ali Taha discussing the emergence of inference engineering as a standalone discipline. Topics include productionizing open models (GLM-5.2, Kimi K3), quantization strategies (including experiments showing 20% throughput gains), speculative decoding, KV-cache movement and compaction, disaggregated prefill/decode, model grafting (adding a vision encoder to a language model), GPU/kernel trade-offs, and the systems-level race (NVIDIA, Rubin, Dynamo) to make frontier models faster and cheaper to serve. The discussion covers implications for infrastructure, continual learning, and video/audio generation workloads.
Read original sourceIVP Seeks $1.8B Fund, Reports 31.1% Net IRR
Institutional Venture Partners (IVP) is raising $1.8 billion for its 19th flagship fund and tells prospective limited partners it has generated a 31.1% net internal rate of return since its 1980 founding. Fund documents reviewed by Newcomer show IVP is seeking premium economics — 2% management fee in year one, 2.25% thereafter, 25% carry stepping to 30% after a 2.5x hurdle, and a 3% GP commitment. The firm’s historical performance includes a 1996 fund that returned 6.7x net DPI (94.5% IRR); more recent vintages range from 1.2x net TVPI (2024) to 2.0x DPI (2015 vintage, 22.7% IRR). IVP highlights investments such as Perplexity, Baseten, ClickHouse, Chainguard, and Abridge, notes a major but not largest Anthropic stake, and appears to have missed OpenAI and SpaceX.
Read original sourceLast Week in AI #250: Frontier AI, GPT-5.6, GLM-5.2
This newsletter/podcast episode (recorded 2026-06-27, published 2026-07-21) summarizes major AI developments: U.S. government gating of frontier models (Anthropic's Mythos-5 approval to limited parties; pressure on Meta for voluntary model review), OpenAI's restricted rollout of GPT-5.6 “Sol”, compute and chip supply-chain moves (OpenAI's Jalapeño ASIC with Broadcom/TSMC, Amazon exploring Trainium sales, Micron investment in Anthropic, Groq fundraising), and open-source progress (GLM-5.2 delivering strong long-context coding performance). The episode also touches on workforce/tax credit initiatives, safety/control roadmaps from research organizations, and broader societal responses to AI.
Read original sourceBaseten: Frequently Asked Questions
What is Baseten?
Baseten is a B2B platform for deploying, serving and managing machine learning models and AI inference workloads in production.
Who uses Baseten?
Its users are developers, machine learning engineers, AI platform teams, data scientists and enterprises building or operating AI applications.
How does Baseten make money?
It earns revenue through usage-based token and compute billing, plus paid platform tiers and enterprise contracts for dedicated or controlled deployments.
Company Facts
- Founded
- 2019
- Core Segment
- B2B SaaS Provider
- Company Size
- 50–200
- Official Link
- baseten.co
