Skip to content
NLEN
Illustration: Local or via an API? The three-year calculation

Local or via an API? The three-year calculation

By Ivo Donker — compiled with AI support (Claude & Gemini) · Last updated: August 7, 2026

Article reviewed on 2026-08-07. Within the canon of gids.llmnet.nl, this article falls under Pillar 2 ("Hardware, performance & energy") and bridges the canon gap regarding the local-versus-API cost comparison over a multi-year horizon. Within the reader's roadmap — choose, install, use, integrate, manage, and serve — this page explicitly positions itself in the first phase: selecting the most suitable delivery model. For the minimum required hardware baseline, consult the article on the minimum hardware baseline for local LLMs to determine which physical specifications are required for your target model size. Whereas previous articles on energy focused on calculating or measuring incidental power consumption, this article provides the comprehensive assessment over a three-year period (36 months). A three-year timeframe was chosen because physical hardware is depreciated both fiscally and technically within that window, while API subscriptions and token costs act as ongoing operational expenditures. Choosing between an in-house local server and an external API service is not a matter of dogma, but a business and technical calculation with specific prerequisites, risks, and trade-offs on both sides.

The building blocks of the three-year TCO calculation

The Total Cost of Ownership (TCO) of an LLM infrastructure over 36 months rests on four fundamental building blocks. To ensure an apples-to-apples comparison, the costs for both the local route and the API route must be evaluated against the exact same benchmarks.

In addition to these four building blocks, the choice of a specific model directly impacts VRAM usage and processing capacity. Without redefining the core technology, the memory footprint can be significantly reduced using quantization. See the explanation of quantization techniques. Read this guide for a detailed breakdown of how weight reduction affects VRAM consumption. Likewise, context length impacts the required compute power. See the guide to optimizing the context window. Read this article for strategies to keep memory usage with large prompts within the limits of your hardware.

Three worked calculation examples over three years

To clarify the financial dynamics, three scenario models have been developed over a 36-month period: heavy use (continuous/production), medium use (daily office tasks), and light use (experimentation and occasional use). All amounts mentioned in this article are assumptions within a calculation model prepared on 2026-08-07.

In the heavy scenario , a continuous load is assumed with a monthly volume of 100 million input tokens and 20 million output tokens. Local hardware here requires a rack server with multiple high-end GPUs. In the medium scenario , an organization processes 10 million input tokens and 2 million output tokens per month on a single powerful workstation GPU. In the light scenario , it concerns 500,000 input tokens and 100,000 output tokens per month, running on an entry-level graphics card or a compact studio computer.

The table below provides a complete overview of estimated costs over three years (36 months). All amounts are assumptions within a calculation model, dated 2026-08-07.

Cost Item (Assumptions 2026-08-07) Heavy use (Local) Heavy use (API) Medium (Local) Medium (API) Light use (Local) Light use (API)
Hardware investment (one-off) € 8.500,00 € 0,00 € 2.800,00 € 0,00 € 1.200,00 € 0,00
Power consumption (per year) € 1.150,00 € 0,00 € 180,00 € 0,00 € 35,00 € 0,00
Maintenance & administration (per year) € 1.200,00 € 300,00 € 450,00 € 150,00 € 150,00 € 50,00
API token costs (per year) € 0,00 € 22.800,00 € 0,00 € 2.280,00 € 0,00 € 114,00
Total power (36 months) € 3.450,00 € 0,00 € 540,00 € 0,00 € 105,00 € 0,00
Total administration (36 months) € 3.600,00 € 900,00 € 1.350,00 € 450,00 € 450,00 € 150,00
Total API costs (36 months) € 0,00 € 68.400,00 € 0,00 € 6.840,00 € 0,00 € 342,00
Total TCO over 3 years (36 mo) € 15.550,00 € 69.300,00 € 4.690,00 € 7.290,00 € 1.755,00 € 492,00

The results show a clear tipping point. Under heavy usage, a breakeven point emerges where the local hardware investment is quickly recouped compared to high-volume API consumption. Under light usage, however, the fixed costs of local hardware remain far too high compared to the low variable consumption of an API.

Hidden and indirect costs of the API route

Although the API route offers the advantage of not requiring upfront capital investments, it introduces specific indirect costs and operational friction that are not immediately evident in the base rates per token.

First, limits on the number of requests per minute (rate limits) can slow down business processes or necessitate maintaining multiple accounts and fallback mechanisms. See the overview of API rate limits and costs. Consult this page to understand how request volume limits directly affect your operational costs. Being forced to switch to more expensive endpoint tiers to achieve higher throughput directly increases the cost per request.

Second, different API providers use varying methods to split text into tokens. A 1,000-word prompt may result in 1,300 tokens with Provider A and 1,550 tokens with Provider B. To financially normalize this discrepancy, standardization is required. See the guide on cross-provider token usage normalization. Read this guide to learn how to compare token counts across different providers in a uniform way.

Third, organizations must take proactive measures against cost overruns. Without tightly configured limits, infinite loops in software can generate thousands of euros in usage within hours. View the guide for the article on monitoring API costs. View this overview for practical methods to prevent unexpected billing spikes with cloud providers.

Finally, cloud providers frequently adjust their rates and discount structures for volume or context caching. An overview of these variable pricing models can be found in the per-token pricing model analysis. Consult this analysis to understand the differences between tiered and flat-rate token pricing.

Weaknesses of the API route

Hidden and indirect costs of the local route

The local route eliminates monthly token invoices, but it introduces fixed management overhead and physical risks that must be carefully weighed.

Maintenance and software updates represent a recurring time investment. A local model does not run in a vacuum: it requires maintenance of the inference engine (such as Ollama, vLLM, or llama.cpp), the operating system, network security, and drivers. When a new model is released, administrators must manually test whether the existing hardware architecture provides sufficient VRAM and support.

Additionally, an on-premises setup does not automatically deliver the same quality or speed as a commercial API. Organizations must conduct their own benchmarks to determine whether the inference speed and output precision meet their requirements. Consult the article on benchmarking quality against cost. Read this article to weigh the output quality of local models against the cost of external APIs.

Accurate measurement is essential to express the actual performance of local systems in tokens per second. Check out the methodology for measuring inference speed. Refer to this guide for standardized protocols to reliably measure the number of generated tokens per second. Only once the speed and hardware depreciation are known can the cost per specific task be determined. For this, see the calculation of the cost per task. Consult this study to calculate the financial cost per specific processing task.

Drawbacks of the on-premises route

The dynamics of a three-year term: scale, obsolescence, and price drops

Over a three-year horizon, the parameters of the equation do not remain static. Two opposing forces affect the TCO analysis over the 36-month cycle.

On the cloud side, there is a historical trend of falling API prices per million tokens. New, more efficient models and increased competition among API providers often cause the cost per token for equivalent performance to decline after 12 to 24 months. Those who exclusively use an API automatically benefit from these price reductions and model upgrades without making new investments.

On the local side, hardware depreciation is fixed. Once purchased, a server remains physically present after three years and is fully depreciated by the end of that period. The marginal cost of processing extra tokens at that point is virtually zero (electricity only). However, physical equipment also ages functionally. A GPU considered modern as of 2026-08-07 may struggle in 2028 with the latest model architectures requiring larger contexts or different quantization formats. This computational uncertainty is a fundamental component of the multi-year evaluation.

Concrete decision rules: when to choose on-premises, API, or hybrid?

Based on the three-year cost structure and operational characteristics, the choices can be summarized into clear decision rules:

  1. Opt for the on-premises route when:
    • Token volume is high and continuous (such as with automated batch processing or 24/7 production systems).
    • The workload is predictable and fits within the capacity of the purchased hardware.
    • Policy requires that data never leaves the internal network under any circumstances.
    • Internal capacity is available for server hardware management and maintenance.
  2. Choose the API route when:
    • Usage varies widely, is intermittent, or has low monthly volumes.
    • You want immediate access to the most advanced commercial models without upfront investments.
    • There is no internal capacity or willingness to manage and maintain server hardware.
    • The application requires fast development cycles with minimal infrastructure complexity.
  3. Choose a hybrid model when:
    • The baseline load is handled locally on on-premise hardware, while unexpected peak loads are routed to the cloud via an API.
    • Sensitive data is processed locally, while non-sensitive tasks are sent to external APIs. To set up such an architecture, read the guide on hybrid cloud-edge LLM integrations. Read this guide to discover how to combine a local baseline with cloud capacity for peak periods.

Privacy and data sovereignty as a cost factor

Privacy is not an abstract legal requirement, but a direct factor in the financial calculation. When using external API services, prompts, documents, and personal data leave your own physical hardware. This requires data processing agreements, compliance with privacy legislation (such as the GDPR), and monitoring to ensure data is not used to train the provider's proprietary models.

With a local setup, not a single byte leaves the local device or the internal network. This eliminates legal mitigation processes and compliance documentation. To assess how privacy requirements impact architectural choices, we refer to the article on privacy-friendly AI architectures. Consult this page for guidelines on setting up data flows that remain entirely within your own network boundaries.

Weaknesses and limitations of this article

This article provides a methodical framework for making a three-year cost comparison, but has the following explicit limitations:

The final decision for your organization depends on your own, pre-measured usage profile and available management capacity. This overview was reviewed on 2026-08-07.