Skip to main content

The Real Price of a Token: Why Sustainable Corporate AI Is Moving On-Premise

Why embedding on-premise AI is the smart move for your business. Every business embedding Artificial Intelligence (AI) into its operations is making a bet on the price of a token. Most do not realise it. Token prices have fallen relentlessly for two years, and boards have built business cases on the assumption that they will stay low. The economics underneath tell a different story. This post explains how token pricing works, where the money behind it comes from, and why the sustainable path for corporate AI is private deployment of open-source models.

How token pricing works

Large language models process text as tokens, fragments of roughly three-quarters of a word. When you use a cloud AI service via an Application Programming Interface (API), you pay per million tokens, with separate rates for input (what you send the model) and output (what it generates). Output tokens cost more because generating text requires more compute than reading it.

Providers set these prices against the cost of inference: the graphics processors serving the model, the electricity and cooling behind them, and the data centre capacity they occupy. As at early 2026, indicative pricing sits at USD$1.25 per million input tokens and USD$10 per million output tokens for OpenAI’s GPT-5, USD$3 and USD$15 for Anthropic’s Claude Sonnet, and as low as USD$0.27 and USD$1.10 for the open-weight DeepSeek V3.

The trend has been one-directional. Epoch AI’s analysis of inference pricing found that the cost to achieve a given level of model performance has fallen between 9 times and 900 times per year depending on the task. DeepSeek V3 launched in late 2024 at roughly one hundredth of GPT-4’s launch price for comparable capability. Two years of this has trained the market to treat intelligence as nearly free.

What it cost to build

The models did not come cheap. Sam Altman confirmed GPT-4’s training cost exceeded USD 100 million, and Stanford’s AI Index estimated Google’s Gemini Ultra at USD 191 million. Estimates for the 2026 frontier class run to USD 200 to 500 million per training run.

Training is the small line item. The infrastructure is the story. Microsoft, Alphabet, Amazon, Meta and Oracle have collectively committed roughly USD 660 to 690 billion in capital expenditure for 2026, nearly double 2025 levels, with about three-quarters directed at AI compute, storage and network. Goldman Sachs projects USD 5.3 trillion in combined capex for the four largest hyperscalers between 2025 and 2030, and the sector raised over USD 100 billion in debt in 2025 alone to fund it.

The revenue on the other side of the ledger

Now compare the return. OpenAI’s revenue run rate reached roughly USD 25 billion in mid 2026, yet the company projects USD 14 billion in losses for 2026 and reportedly expects to spend well over USD 100 billion on compute commitments. Anthropic’s reported annualised revenue passed USD 30 billion, with compute costs consuming an estimated 60 per cent of it. Industry analysis is blunt: the major providers are pricing inference below the cost of serving it, funded by venture capital and debt, to build market share and dependence.

That is the definition of a subsidy. Tokens are cheap because investors are paying part of your bill. Every dollar of that USD 5.3 trillion infrastructure programme expects a return, and the arithmetic only resolves two ways: usage grows into the capacity at sustainable margins, or prices rise. When investors stop funding losses and start demanding return on investment, the organisations most exposed will be those that wired subsidised tokens into their core processes and shed the capability to work any other way.

The sustainable alternative: own the model

None of this is an argument against embedding AI. Agentic processes and embedded models are becoming as fundamental to the workplace as email, and organisations that delay will fall behind. It is an argument about how you deploy.

Open-source and open-weight models now match or approach frontier performance for most business workloads, and they can be deployed privately, on dedicated hardware, at a cost you control. NVIDIA’s DGX Spark illustrates how far the economics have shifted: a desktop unit with 128 GB of unified memory and a petaFLOP of AI compute, capable of running models up to 200 billion parameters. Thats why embedding on-premise AI is the smart move for your business. It’s a dramatic shift from a metered dependency on someone else’s balance sheet, with the added benefit that your data never leaves your environment.

Unisphere Solutions’ on-premise AI offering builds on exactly this approach: open-source models deployed privately on NVIDIA DGX-class hardware, sized to your workloads, integrated with your systems and governed under your policies. All managed, secured and updated for you by experts. You get the productivity of embedded AI and agentic workflows without exposure to a pricing model that is one investor meeting away from repricing.

Token prices will not stay subsidised forever. Boards that want AI in the business for the long term should own the capability, not rent it. Talk to Unisphere Solutions at unisphere.co.nz about deploying AI privately and sustainably.

Discover more from Unisphere Solutions

Subscribe now to keep reading and get access to the full archive.

Continue reading