Inference API
Simple, Transparent Pricing
Pay-as-you-go rates for the Inference API. No minimums, no commitment - pay only for the tokens you use.
- EUR Native Pricing
- No Minimums, No Commitment
- EU Sovereign by Default
Inference API pricing
Pay-per-token rates for our shared, EU sovereign inference platform. Need reserved capacity or your own hardware? Those two are priced per engagement.
You're here
Inference API
Shared infrastructure, pay-per-token. Start in minutes, scale as you grow - no minimums, no commitment.
- EU sovereign by default
- Pay as you go
- OpenAI-compatible API
How API Pricing Works
The Inference API is priced per token. A token is roughly 4 characters in English (this varies by model and language). You pay for what you use - no minimums, no commitments.
What is a token?
Tokens are the basic units LLMs process. In English, 1 token ≈ 4 characters or ¾ of a word. 1,000 words ≈ 1,300 tokens. Other languages may use more tokens per character.
Input vs Output
You pay separately for input (your prompt) and output (the model's response). Output costs more than input because your prompt is read in a single parallel pass, while each output token needs its own full pass through the model - one token at a time.
Choosing a model
Each model excels at different tasks - there's no single "best" choice. EU Sovereign models (🇪🇺) guarantee your data stays in European jurisdiction.
EU
EU Sovereign Models
Full GDPR compliance, no US CLOUD Act exposure
| Model | Modality | Input/1M | Output/1M | Context |
|---|---|---|---|---|
| MiniMax M2.7 Ultraspeed | Text | €0.60 | €2.40 | 192K |
| Gemma 4 31B | Text + vision | €0.20 | €0.35 | 128K |
| gpt-oss-120b | Text | €0.22 | €0.59 | 128K |
| Whisper Large v3 | Speech-to-text | €0.10 | - | - |
| E5-Mistral 7B | Embeddings | €0.13 | - | 4K |
Prices in EUR, excl. VAT. EU-hosted models include full data sovereignty. Speech-to-text and embeddings are billed on input tokens only (they return transcripts or vectors, not generated text).
Global
Global Model Catalog
Additional models via global infrastructure
| Model | Modality | Input/1M | Output/1M | Context |
|---|---|---|---|---|
| Llama 3.3 70B | Text | €0.60 | €1.20 | 128K |
| DeepSeek V3.1 | Text | €3.00 | €4.50 | 128K |
| DeepSeek V3.2 | Text | €3.00 | €4.50 | 32K |
Prices in EUR, excl. VAT. Requests processed on global infrastructure outside the EU.
Estimate your monthly cost
Start from a typical workload, then adjust the numbers to match yours.
Estimated monthly cost (excl. VAT)
€324.00
- Per request
- €0.00108
- Per day
- €10.80
- Rate (in / out per 1M)
- €0.60 / €2.40
EU Sovereign - data stays in EU
Estimate only. Assumes ~30 days of steady usage. Actual token counts vary by language, content, and model tokenizer (~750 words ≈ 1,000 tokens for English). Speech-to-text and embeddings aren't shown here - they bill on input tokens only; see the rate table above. Check exact usage anytime at cloud.infercom.ai/plans/usage.
Access Tiers
Choose how you want to start - upgrade anytime as you grow.
Start here
Developer
Pay as you go
Sign up, create an API key, and pay only for the tokens you use. No minimums, no commitment.
Enterprise
List price, with volume discounts
The same per-token rates as your baseline, with committed-volume discounts, custom SLAs, priority support, and higher rate limits.
Frequently Asked Questions
How does billing work?
The Inference API is pay-as-you-go. You add a billing method - a credit card - and usage is drawn down per token against the published list prices, so you only ever pay for what you use, with no minimums and no monthly commitment. Cost is easy to model up front from the rates above.
Can I see my token usage and current spend?
Yes, in real time. Your live token consumption is at cloud.infercom.ai/plans/usage, and your current billing and invoices are at cloud.infercom.ai/plans/billing. You always know exactly what you've used and what it costs.
What are the rate limits?
Rate limits vary by plan and model. The Developer tier has standard rate limits documented in our rate limits documentation. Enterprise plans offer custom rate limits tailored to your workload. Rate limits control request frequency - they do not affect inference speed per request.
What payment methods do you accept?
For our inference service, we accept major credit cards through Stripe. Enterprise, dedicated capacity, and on-premises customers can also pay by invoice.
What's included in EU sovereignty?
For EU sovereign models, all data processing happens in our European data centers by default. Your data never leaves EU jurisdiction unless you explicitly opt to use Global Catalog models. Full GDPR compliance and AI Act readiness included.
Do you offer volume discounts?
Yes - both Enterprise and dedicated capacity contracts offer custom pricing based on your committed usage. Contact our sales team to discuss your requirements.