VL

vLLM

Open-source LLM inference and serving engine.

Available information varies by company and source.

Profile record updated:

Company facts

Entity type
COMPANY
Company size
50–200
Market role
B2B SaaS Provider
Official website
vllm.ai

What vLLM does

vLLM operates as an open-source infrastructure project for LLM inference and serving. It creates value by reducing the performance and memory costs of running large models, which supports adoption among developers, research groups, and commercial AI teams. Distribution is product-led through the website, documentation, GitHub repository, and community channels. The provided materials establish ecosystem adoption and collaboration interest, but they do not disclose a formal paid software packaging model.

Category differentiation

vLLM is not a foundational model provider and it is not a consumer AI app. It is inference and serving infrastructure used to run large language models efficiently.

Strategic context

AI-supported assessment from the existing company research; distinguish interpretation from sourced facts.

vLLM is an open-source software project that provides a high-throughput, memory-efficient inference and serving engine for large language models. It is positioned as infrastructure for organisations and developers that need to deploy, serve, and operate LLM workloads in production rather than as a consumer application or advertising product. The project creates value by improving serving efficiency for LLM deployments and by building a contributor and user ecosystem around its open-source runtime. Its direct users are AI engineers, platform teams, researchers, and companies building LLM-based products or services. The provided materials do not disclose a formal legal entity name, a verified headquarters location, or an explicit commercial pricing model.

Company news briefing

Briefing updated:

vLLM continues to reinforce its role as a core inference serving engine, expanding its cross-hardware integration across NVIDIA Blackwell platforms, Grace Blackwell systems, and emerging Google TPUv7 Ironwood stacks via the TorchTPU framework. Recent developments also highlight system-wide optimizations such as vLLM-Omni for scaling multimodal stacks like MiniMax H3, alongside ongoing optimizations for PagedAttention and native support for new model architectures like DiffusionGemma.

Business model & monetisation

No explicit monetisation model is disclosed in the provided materials. The project is distributed as open-source software and signals collaboration with companies using vLLM in products or services, but no formal subscription, licensing, usage-based pricing, or support plan is stated in the input.

Open-source software distribution
Commercial collaboration with companies using the software

Products & capabilities

No products with linked sources are available in this view.

Recent recorded signals

Dates refer to the source publication. Older entries are historical context, not evidence of a new event.

  • MiniMax H3 on vLLM-Omni: From System-Wide Optimization to Real-Time Serving with FastVideo’s FastH3

    vllm.ai

    Recorded impact score: 5/5

    How vLLM-Omni optimizes and scales the complete MiniMax H3 stack, then integrates FastVideo’s four-step FastH3 for generation faster than playback.

  • Qwen3-8B inference benchmark and FP8 on Blackwell

    dev.to

    Large Language Models (LLM) & AI · Recorded impact score: 2/5

    Independent benchmarks compare Qwen3-8B inference on an RTX PRO 6000 Blackwell (96 GB) across three serving stacks (vLLM 0.27.1, SGLang 0.5.9, and llama.cpp CUDA). At concurrency 32 using BF16, vLLM achieved 1,725 aggregate tokens/s (TTFT p50 39 ms), SGLang 1,327 tok/s (TTFT p50 42 ms), and llama.cpp 428 tok/s (TTFT p50 316 ms). Applying an FP8 checkpoint to vLLM increased throughput by ~1.5x (aggregate 1,725 -> 2,597 tok/s; single-stream 86 -> 130 tok/s) with lower latency and no detected regressions on a fixed factual check. The author documents methodology, reproductions, and an sm_120-specific kernel workaround required to run FP8 on workstation Blackwell hardware.

    • GPU used: RTX PRO 6000 Blackwell, 96 GB (workstation Blackwell, sm_120).
    • Model benchmarked: Qwen3-8B across vLLM 0.27.1, SGLang 0.5.9, and llama.cpp (CUDA).
  • Tokens-per-Second Benchmarks Explained

    dev.to

    Large Language Models (LLM) & AI · Recorded impact score: 2/5

    This technical guide explains what "tokens per second" (tok/s) actually measures for local LLM inference, why single-user tok/s numbers can be misleading, and how concurrency, batching, and prompt processing change the observed speed. It contrasts single-user latency with server throughput, highlights vLLM's continuous-batching advantage versus Ollama under high concurrency, defines related metrics (P99 latency, time to first token / TTFT), and provides practical measurement advice using tools like Ollama and vLLM and calculators from notAcalculator. The article also gives realistic tok/s expectations for different model sizes on consumer hardware and lists practical tips for reading and running benchmarks yourself.

    • Tokens are the unit of both billing and speed for LLMs; tokenization affects cost and measured tok/s.
    • Under a Red Hat benchmark on an A100 40GB with Llama 3.1 8B, vLLM peaked around 793 tok/s combined throughput versus about 41 tok/s for Ollama at high concurrency (~19x gap).
  • One TPU Chip, Eight Agents: Serving Small Agent Workloads

    dev.to

    Large Language Models (LLM) & AI · Recorded impact score: 2/5

    An engineer implemented a pure-JAX serving path to run a Gemma 4 E2B quantization-aware-trained (QAT) checkpoint on a single Cloud TPU v6e chip because vLLM could not load the QAT export on TPU. The author created a safetensors→JAX loader and a JAX decode kernel, validated correctness against full re-forward, and measured kernel decode rates up to ~2,888 tok/s (int4/int8 donated path). Memory math shows eight 8K contexts fit comfortably on a 32 GB HBM v6e chip (≈1.21 GB KV for eight 8K contexts). However, the experimental server lacks prefix caching, guided/schema-constrained decoding, and continuous batching, so end-to-end HTTP serving without batching reached only ~139–143 aggregate tok/s with latency rising under contention. Verdict: viable experimental path for cases that need the QAT checkpoint, but not yet a drop-in vLLM production replacement.

    • Author implemented a pure-JAX Gemma 4 E2B loader and serving engine (safetensors → JAX PyTree) to run a QAT checkpoint on TPU.
    • vLLM could not load the QAT Gemma 4 E2B exports on TPU due to loader assumptions about K/V-side tensors (filed as tpu-inference #3225).

Explore company relationships

Questions about vLLM

What is vLLM?

vLLM is an open-source inference and serving engine built to run large language models with high throughput and memory efficiency.

Who uses vLLM?

vLLM is used by developers, AI engineers, research teams, and enterprise platform teams that deploy and serve LLM workloads.

How does vLLM make money?

The provided materials do not disclose a formal monetisation model. The project is distributed as open-source software and signals collaboration with companies using it commercially.

Sources & coverage

This profile uses public, official and technically observable information. Missing information does not prove that a product or relationship does not exist. The list below does not imply that every profile statement has been verified.

8 publicly documented primary sources and citations linked across the market graph.

Continue your research on vLLM

Explorer includes additional company details, a Watchlist for up to 25 companies and your personal Strategic Intelligence Agent. It monitors your market daily and delivers tailored briefings with clear strategic context whenever relevant news occurs.

Free, with no time limit.