Vivold Consulting

A 14-person open-source team just became the default way 8.9M developers run local AI - and a lever for slashing inference bills

Key Insights

Ollama, the open-source tool that lets developers run open-weight AI models on their own machines in minutes, raised a $65M Series B led by Theory Ventures ($88M total), revealing it now serves 8.9 million developers monthly and sits inside 85% of the Fortune 500 - with just 14 employees. Founders Jeff Morgan and Michael Chiang previously built Docker Desktop, and they're repeating the play: abstract away the hardware pain, then monetise a cloud tier priced on GPU time rather than tokens. The backdrop is the industry's loudest cost debate: every company with heavy inference bills is under existential pressure to shift routine workloads to open models.

Stay Updated

Get the latest insights delivered to your inbox

The quiet infrastructure story behind the AI cost debate

While the trillion-dollar IPOs grab headlines, one of the most consequential funding rounds of the week went to a 14-person company. Ollama raised a $65 million Series B led by Theory Ventures, following a $15 million Series A led by Benchmark's Peter Fenton, for $88 million raised in total. The traction numbers explain the investor enthusiasm: launched in 2023 to make open-weight models runnable on an ordinary PC in minutes, Ollama now counts 8.9 million monthly developers, presence in 85% of the Fortune 500, and 176,000 GitHub stars with nearly 17,000 forks.

Docker's playbook, replayed for AI

Founders Jeff Morgan and Michael Chiang know exactly what product ubiquity among developers looks like - they helped build Docker Desktop after Docker acquired their startup Kitematic. The analogy is almost literal: Docker abstracted away hardware configuration so cloud apps could run anywhere; Ollama does the same for open models, which in 2023 were built for researchers and painful for ordinary programmers to stand up. The monetisation layer follows the same freemium logic: the free desktop tool remains untouched, while a cloud service hosts larger models under subscription tiers from free to $100 per month - notably metered on GPU time rather than token limits, a pricing model worth watching as buyers grow weary of opaque token math.

Why now: the open-model inflection

Morgan pins the business inflection to January, when agentic assistants took off and large open models suddenly proved capable of real work like coding. That fuels the industry's sharpest cost argument: Fenton frames it as no longer either/or between open and closed models, but says any company with high inference expenses now has a vital, existential project to shift workloads toward open weights. Ollama is one of a whole crop of open-source projects turning into venture-backed companies - inference providers like Inferact (vLLM) and RadixArk (SGLang), assistant alternatives like NanoClaw, and small model builders like Arcee. There is friction, too: parts of the community accused Ollama of drifting toward commercialisation a year ago, criticism the founders answer by insisting the free local product is unchanged.

Where the money is for you

- The practical takeaway is a hybrid model strategy: route routine, high-volume workloads (classification, extraction, internal chat, first-pass drafting) to open models served locally or via commodity inference, and reserve frontier closed models for the tasks that genuinely need them. Companies making this split are the reason Ollama sits in 85% of the Fortune 500.
- Local execution is also a privacy and compliance play: data that never leaves the machine sidesteps a whole category of vendor-risk review, which is often what unblocks AI pilots in regulated teams.
- If you advise clients, an inference-cost audit is now a legitimate standalone engagement: map which workloads run on premium closed models today, benchmark open-weight equivalents, and quantify the delta - the existential framing from Ollama's own board suggests the savings are frequently material.
- One caution for procurement: a 14-person vendor is a thin organisation to bet critical infrastructure on, however impressive the penetration. The mitigation is the product's open-source core - your exit path is the code itself, which is precisely why this model of company keeps winning enterprise trust.

Related Articles

Discovery Loop aims to automate science itself - and Google is funding the startup draining its own bench, as Hassabis exits the DeepMind CEO role

Jeff Dean, Google's chief scientist and 30th employee, is leaving after 27 years to found Discovery Loop, a public benefit corporation using AI to automate scientific research - taking co-founders Sanjay Ghemawat, Quoc Le (Google Brain), and Oriol Vinyals (DeepMind) with him. Google is a founding investor and cloud partner, supplying compute for at least the first year, with Radical Ventures and Khosla Ventures co-leading the seed. In the same announcement, Demis Hassabis steps down as DeepMind CEO to become chairman and Alphabet chief scientist, with Koray Kavukcuoglu taking over Gemini model development. Alphabet stock fell about 4%.

Abbott orders audits of every new project as ERCOT's queue hits 474GW, roughly 90% of it data centres

Governor Greg Abbott announced that all new Texas data-centre projects must be audited by the Public Utility Commission and grid operator ERCOT - a sharp turn for a state whose loose regulation and cheap power made it second only to Virginia for data centres. The trigger is a staggering queue: ERCOT's interconnection requests doubled from 233GW in January to 474GW, about 90% data centres, more than five times the grid's all-time peak demand. Audits will demand power and water use, noise mitigation, light controls, tax-incentive use, and ownership details - after a voluntary survey that most operators simply ignored.

Volta and Bitdeer will build a 133MW Nvidia Vera Rubin data centre in Norway - Anthropic's latest move in a compute land grab

Anthropic has reportedly signed a $10 billion, six-year compute deal with Volta, an AI cloud startup founded only earlier this year, per Bloomberg. Volta is partnering with crypto-mining firm Bitdeer to develop the data centre - located in Norway, delivering 133 megawatts, and running Nvidia's Vera Rubin architecture - and is a member of Nvidia's Cloud Partner programme. It caps an aggressive capacity spree that also includes recent compute deals with SpaceX and Amazon, as Anthropic races rivals for the scarcest input in the industry.