For two years the case against open-weight models was simple: they were not good enough, and the licenses were murky. Both objections have mostly dissolved, which makes the remaining question honest at last. It is a money question, and the money question has a sharper answer than the discourse suggests. At mid-2026 prices, the cheap way to run an open model is almost never to own the machine it runs on. Owning pays only under conditions most small operations never meet, and knowing the conditions is worth more than holding an opinion about openness.
The gap closed. The math did not.
Start with capability, measured two ways. Artificial Analysis put the best open-weight models 6 points behind GPT-5.5 on its composite intelligence index in April 2026; a year earlier the gap was roughly 13. Stanford’s 2026 AI Index, reading Arena preference scores instead, has the top closed model 3.3 percent ahead of the top open one as of March 2026, slightly wider than the near-parity of August 2024, with six of the top ten models closed. Pick either instrument and the conclusion is the same shape: the frontier is still closed, and the distance to it is now single digits, close enough that for bounded business tasks the difference often disappears inside your evaluation noise.
The licensing fog lifted too. DeepSeek’s current open flagship ships under MIT. OpenAI’s gpt-oss and the whole Mistral 3 family carry Apache 2.0. Even Llama’s bespoke community license only bites past 700 million monthly active users, which is not your problem. For a deploying operator, the lawyers are no longer the blocker.
So the blocker is arithmetic, and the arithmetic has three tiers.
”Going open” is three different purchases
The first tier is hosted open-model tokens: someone else’s GPUs, your choice of open weights. On Together’s July 2026 price list, gpt-oss-120B costs $0.15 per million input tokens and $0.60 out, and DeepSeek V4 Pro runs $1.74 and $3.48. Set those against the closed flagships (GPT-5.5 at $5 and $30, Claude Opus 4.8 at $5 and $25) and the discount on output tokens runs roughly seven to ten times. Together sells open-model inference, so read its list with that in mind, but its DeepSeek prices match its competitors’ within cents.
The second tier is where the arbitrage quietly collapses: the closed vendors’ small models. GPT-5.4-nano lists at $0.20 and $1.25. Claude Haiku 4.5 at $1 and $5. Claude Sonnet 5, an agent-grade model, at an introductory $2 and $10 through August 31, 2026. These sit in the same band as hosted open weights, and they arrive with no serving decisions at all. (One caution when comparing across vendors: per-token prices assume comparable tokenizers, and Anthropic’s own docs note its newest models produce roughly 30 percent more tokens for the same text. Compare cost per task, never cost per token.)
Which means the honest comparison was never open versus closed. It is the cheapest model that clears your evaluation, from any source, the discipline the small-models field notes argued from the build side. Open weights won that contest often enough to belong in every bake-off. That is a different claim than “run your own GPUs.”
The tier where the bill grows teeth
The third tier is true self-hosting, and its floor is visible on any price sheet: an H100 rents for $2.89 to $4.29 an hour across the major GPU clouds, call it $2,100 to $3,100 a month for one always-on card before redundancy, before storage, before anyone is paid to keep it alive.
The fuller accounting comes from a practitioner analysis published in May 2026 by Azumo, a services firm that builds exactly these deployments, so treat its numbers as a vendor’s field estimate rather than a study. Standing up production self-hosting: two to four engineers for three to six months, $200,000 to $300,000. Keeping it running: $125,000 to $150,000 a year in on-call engineering. Teams plan for 70 to 80 percent GPU utilization and achieve 20 to 40. At a million-per-day scale (a billion tokens a month), the same analysis prices the API path at $1,000 to $2,000 a month against $22,000 to $31,000 self-hosted. Its break-even lands near 10 billion tokens a month, sustained, which is two orders of magnitude past most SMB workloads. Even the survey most friendly to buying, a16z’s January 2026 poll of 100 enterprise executives, concedes only that total cost of ownership between open and closed is “converging,” not that ownership wins.
And the target moves. Epoch AI measured the price of reaching a fixed capability level falling by a median of roughly 50 times per year through early 2025, and promotional prices this summer expire on dates you can put in a calendar. A break-even computed today decays while your procurement cycle runs.
What the market’s revealed preference says
Behavior corroborates the math. In Menlo Ventures’ survey of 495 enterprise decision-makers, open models’ share of enterprise LLM usage fell from 19 percent to 11 percent in the year the capability gap narrowed most, while the a16z survey found preference for closed models still rising. The counter-current is real but specific: cost-sensitive developer traffic keeps shifting open, and the strongest named case runs through a router, not a rack. Airbnb’s CEO says the company leans heavily on Alibaba’s Qwen for customer-service work because it is fast and cheap, inside a portfolio of thirteen models (his remarks reached print via the Alibaba-owned SCMP, worth knowing when you weigh them). Nobody in that story bought a data centre. They bought the cheapest tokens that passed their tests.
The four conditions where owning wins
Self-hosting stops being a false economy when specific conditions hold, and they compound rather than substitute:
- Volume that survives honesty. Sustained work in the billions of tokens per month after you apply real utilization rates, not planned ones. Near the 10-billion mark, the pure compute math genuinely flips.
- A narrow task a small model can own. Fine-tuning a model under 16B parameters costs $0.48 per million training tokens on Together’s current list; a 10-million-token tune is single dollars of compute. The expensive part is the evaluation set proving it beats the API alternative, which you need anyway.
- Obligations that price control differently. If your sector or your contracts demand data residency and provable custody, the comparison was never purely per-token. Canada is spending $2 billion on sovereign compute capacity, which will widen the domestic options; note the strategy funds infrastructure, it does not impose a residency duty by itself. Your duties come from your regulator and your customers.
- The discipline already in place. Monitoring, patching, a model-refresh cadence, and the portability that keeps any vendor swappable. Ownership concentrates operational risk on your own team; it is a reward for maturity, never a shortcut to it.
Meet all four and self-hosting is a defensible line item. Meet three or fewer and the better version of “going open” is the one that needs no capex: put open weights in the bake-off, buy whichever tokens clear your eval cheapest, and re-run the numbers each quarter. Every price in this piece was true on July 3, 2026. The one certainty about them is the direction: down, fast, and not waiting for your budget cycle.