Ollama

Local and cloud infrastructure for open-model AI development.

Available information varies by company and source.

Profile record updated:

Company facts

Official name
Ollama Inc.
Entity type
COMPANY
Founded
2023
Headquarters
United States
Company size
<10
Market role
B2B SaaS Provider
Official website
ollama.com

What Ollama does

Ollama operates a developer-infrastructure model. It distributes a free local runtime to drive broad adoption among developers building with open models, then monetises usage when those workloads need larger hosted models, higher throughput, or managed cloud execution. The company creates value by combining local privacy, offline capability, model management, API compatibility, and cloud scalability in a single workflow, reducing friction for teams moving from experimentation to production.

Category differentiation

Ollama is a developer infrastructure company for running open models locally and in the cloud. It is not a consumer AI chatbot, media company, or advertising technology platform.

Strategic context

AI-supported assessment from the existing company research; distinguish interpretation from sourced facts.

Ollama Inc. is a private US software company that provides developer infrastructure for running and deploying open large language models. Its core product is a local runtime and toolchain that lets developers download, manage, and serve models on-device through a CLI and OpenAI-compatible API. The company also offers Ollama Cloud, a hosted inference layer for larger models and higher-throughput workloads, extending the same workflow from local environments into managed cloud compute. The business serves developers, machine learning engineers, and technical teams building applications that require private, offline, or hybrid AI inference. Ollama creates value by simplifying open-model adoption, reducing setup friction, and enabling a consistent local-to-cloud development path. It makes money through paid cloud subscriptions and infrastructure usage tied to hosted inference capacity, while the free local runtime acts as the primary adoption engine.

Company news briefing

Briefing updated:

Ollama continues to expand its ecosystem presence, securing day-0 support for DeepSeek's V4.1 Flash model alongside its ongoing integration into enterprise frameworks like IAB Tech Lab’s AAMP 2.3. Building on its transparent per-token pricing across Pro, Max, and Team plans, Ollama remains a core component for local development stacks, including Azure Cosmos DB vector search workflows and enterprise Java agent architectures.

Business model & monetisation

Ollama monetises through a hybrid freemium SaaS and pay-per-use infrastructure model. The local runtime is free, which expands developer adoption and community usage. Revenue comes from Ollama Cloud through paid tiers such as Pro and Max, with recurring subscription pricing and usage allowances tied to hosted inference capacity, concurrency, model scale, and compute-intensive workloads.

Hosted cloud subscriptions
Software Subscription
Inference capacity and usage allowances
Pay-per-Use

Products & capabilities

No products with linked sources are available in this view.

Products & market categories

Recent recorded signals

Dates refer to the source publication. Older entries are historical context, not evidence of a new event.

  • DeepSeek Launches V4.1 Flash with Novel Encoder-Decoder Architecture

    latent.space

    AI Model Launch · Recorded impact score: 5/5

    DeepSeek released DeepSeek-V4.1-Flash, a 763B-parameter mixture-of-experts model employing a novel causal encoder-decoder architecture with 8B active parameters for prefill and 16B for decode. It features native vision understanding, 1M token context, an MIT license, and extreme inference efficiency, claiming up to 1/8 KV cache footprint versus V4 Flash. Independent evals (Artificial Analysis Index 40, Vals Index #1 open-weight) show it surpasses V4 Pro at lower cost. API pricing is $0.30/1M input and $1.20/1M output tokens. DeepSeek has soft-retired V4 Pro, routing traffic to V4.1 Flash. The model supports SSD offload and local deployment, with Ollama and Baseten offering day-0 support. Technical discussions highlight the architecture's novelty and potential impact on long-context agents.

    • DeepSeek launched V4.1-Flash with a causal encoder-decoder architecture, 763B total params (8B prefill/16B decode active).
    • Artificial Analysis Index scores V4.1-Flash at 40, above V4 Pro and below GLM-5.3-Flash.
  • Developer Builds Autonomous AI Agent to Hunt Paid Bounties

    dev.to

    Large Language Models (LLM) & AI · Recorded impact score: 2/5

    A developer built an autonomous AI agent that scans hundreds of online gig/bounty listings, filters scams and human-only tasks, generates deliverables using live market data and a local LLM, and notifies a human for approval. The stack uses free tools (Python orchestration, Ollama with a local model, Chart.js, public crypto APIs, GitHub Pages, Windows Task Scheduler) resulting in $0/month infrastructure cost. In 48 hours the agent found many listings but only a handful were actionable due to geo-walls, ghost sponsors, and other filters; the author highlights the need for revenue tracking and human-in-the-loop oversight.

    • Author built an autonomous AI agent that scans 232+ listings across multiple platforms to find paid work and generate deliverables.
    • Stack used: Python orchestration, Ollama with qwen3:4b (local LLM), Chart.js, CoinGecko, DeFiLlama, Solana RPC, GitHub Pages, and Windows Task Scheduler.
  • Ollama's transparent pricing

    ollama.com

    Recorded impact score: 3.5/5

    Ollama's Pro, Max, and Team plans now use industry-standard per-token pricing with usage included on every plan.

  • Flash Onyx 2.2 Released, Finishes Model Improvements

    dev.to

    Large Language Models (LLM) & AI · Recorded impact score: 2/5

    Flash Onyx 2.2, a new version of the Flash Onyx model by the author 'Natuworkguy', is published and available via Ollama. The release ships in two sizes (12b for consumer hardware and 31b for GPU-equipped machines) and focuses on improving request comprehension, producing correct concise answers, automatic stack selection, safer and more accurate Manim animation code generation, and better game-development patterns. Sampling and context settings were adjusted (temperature 0.7→0.6, min_p 0.0→0.05, num_ctx 32768→65536). The post includes install/pull commands for Ollama, recommended MODEL env settings, and an install script link. The author invites feedback and notes that future work (2.3) will address any failures encountered when using 2.2.

    • Flash Onyx 2.2 is published and available on Ollama.
    • The model is offered in two sizes: 12b (runs on normal consumer hardware) and 31b (requires a real GPU).
  • Flash Onyx 2.2: Local Model for Law and Game Feel

    dev.to

    Large Language Models & Local Agents · Recorded impact score: 2/5

    Flash Onyx 2.2 is a development update for a local-first agent shell (Flash Onyx) built around gemma4 configured as an engineering agent that runs fully on a 16 GB laptop GPU at single-digit tokens per second. Version 2.2 expands the system-prompt domains beyond code into legal and game-design guidance, with strict rules (e.g., refuse invented citations, jurisdiction-first for law; game feel heuristics for games). The author compressed prompt text to reduce per-conversation prefill cost, tuned sampling (temperature moved from 0.7 to 0.6 and min_p from 0 to 0.05), and increased context window to 65536 after probes showed 32k/64k/128k used ~8.1–8.2 GB GPU memory. Build tooling now injects the repo license into models; Onyx 2.2 is not yet published but is planned as Natuworkguy/flash-onyx-2.2 in 12b and 31b variants.

    • Flash Onyx 2 is based on gemma4 configured as an engineering agent.
    • The model runs locally on a 16 GB laptop GPU at single-digit tokens per second.

Explore company relationships

Questions about Ollama

What is Ollama?

Ollama is a B2B software platform that lets developers run, manage, and deploy open large language models locally and through a hosted cloud layer.

Who uses Ollama?

Ollama is used by developers, machine learning engineers, startups, and technical teams that need private, local, or scalable hybrid AI inference.

How does Ollama make money?

Ollama makes money through paid cloud plans and hosted inference usage, while its local runtime is offered free to drive developer adoption.

Sources & coverage

This profile uses public, official and technically observable information. Missing information does not prove that a product or relationship does not exist. The list below does not imply that every profile statement has been verified.

17 publicly documented primary sources and citations linked across the market graph.

Continue your research on Ollama

Explorer includes additional company details, a Watchlist for up to 25 companies and your personal Strategic Intelligence Agent. It monitors your market daily and delivers tailored briefings with clear strategic context whenever relevant news occurs.

Free, with no time limit.