Inference API

Fast, OpenAI-compatible inference

with EU sovereign models you can choose

OpenAI-compatible inference on current open-weight models, on purpose-built hardware in Munich. Choose EU sovereign models and your data stays in the EU. Up to 10x faster than GPUs, pay only for what you use, live in minutes.

Why Infercom

Built for serious builders

Fast answers for production workloads: compliant, and yours to integrate in minutes, with no infrastructure to manage.

Trustworthy

  • EU jurisdiction, data center in Munich
  • No US parent, no CLOUD Act exposure
  • No persistent storage by default

Our Trust Center

Seamless

  • OpenAI- and Anthropic-compatible
  • Latest open-weight models, no lock-in
  • Streaming, tools, structured output

Read the docs

Ultra Fast

  • Up to 10x faster than GPUs
  • Hundreds of tokens per second, per request
  • Purpose-built for inference, not training

See the benchmarks

Why Speed Matters

Speed decides what you can build

It shows up two ways: it compounds across an agent's many calls, and you feel it the instant a real-time app responds.

1 · It compounds

For agents, latency multiplies

A chatbot makes one call per turn. An agent makes 50-100 inference calls to complete a single task - each one waiting on the last.

At that scale, per-call latency stops being a detail and starts deciding what you can ship. Same agent, same logic, same model: the infrastructure underneath is the difference between a task that finishes in minutes and one that takes ten.

50-turn agent task

GPU at 90 tok/s~9.8 min
Infercom~2.5 min

Calculated: each call 10K input and 1K output tokens. Infercom at the measured 3.023 s end-to-end on MiniMax M2.7 (July 2026); GPU at 90 tok/s plus the same time to first token. Not a measurement of a GPU provider.

2 · You feel it

For real-time apps, the first token is everything

In a voice assistant or a live coding tool, what makes it feel human isn't total throughput - it's how fast the response starts. A long pause before the first word breaks the illusion of a conversation.

The dataflow architecture delivers a fast time to first token, so real-time apps feel responsive instead of laggy. Every API response reports its measured TTFT, so you can hold your app to it.

Time to first token

  • Voice agentsNo awkward pause before the reply
  • Live coding assistantsSuggestions that keep up with typing
  • Interactive chatAnswers that start streaming instantly

Integration

One line to switch

Point the OpenAI SDK at Infercom and keep the rest of your code. Chat completions, streaming, structured output, and function calling all work the way you expect.

Python
from openai import OpenAI

client = OpenAI(
    base_url="https://api.infercom.ai/v1",
    api_key="your-api-key"
)

response = client.chat.completions.create(
    model="MiniMax-M2.7",
    messages=[
        {"role": "user", "content": "Hello!"}
    ]
)
# usage includes measured tokens/sec + TTFT

EU Sovereign models - in EU data centers

Global Catalog - processed outside the EU

An extended catalog of additional open-weight models, processed outside the EU, in a region we do not control, for teams that need broader model choice.

For GDPR-sensitive data, use the EU Sovereign models. The Global Catalog runs on external infrastructure outside EU jurisdiction.

View full model catalog

Access Tiers

Start building, scale to production

Self-service on shared infrastructure. Move up as your usage grows.

Self-Service

Developer

Pay as you go

Sign up, create an API key, and start building. Pay only for the tokens you use, against published list prices.

Start Building

Committed Use

Enterprise

Custom

For production workloads that need higher throughput, priority, and a commercial agreement.

  • Production rate limits
  • Scheduling priority
  • Volume discounts
  • Dedicated support
Contact Sales

See per-model pricing

FAQ

Questions, answered straight

Is it really OpenAI-compatible - can I point my existing code at it?

Yes. It's a drop-in replacement for the OpenAI API: change the base URL to https://api.infercom.ai/v1, use your Infercom API key, and the rest of your code stays the same. We're also Anthropic-compatible, so the Anthropic SDK works too. See side-by-side examples on the API compatibility page.

How do I keep my data in the EU?

Use the EU Sovereign models. They run on our hardware in our data center in Munich, and your data stays in EU jurisdiction. We also store nothing persistently: we don't store prompts or responses on disk, don't train on your data, and don't serve cached answers. The Global Catalog models are processed outside the EU, so choose EU Sovereign models for GDPR-sensitive workloads. More detail on our Trust Center.

Are the models quantized?

No. Today we serve every model at the precision its maker released it in, without quantizing it.

Can we connect privately, without going over the public internet?

Not today: the API is reached over HTTPS on the public internet. If private connectivity is a requirement, talk to us about your network.

How fast is it, really?

Measured on our production API: 428 tok/s on MiniMax M2.7 and 713 tok/s on gpt-oss-120b (July 2026). Output speed per request holds when many requests run in parallel; the time to first token grows when requests queue at peak times. Every response reports its time to first token and tokens per second, so you can check it on your own workload.

What's the difference between the Developer and Enterprise tiers?

Developer is self-service and pay-as-you-go with standard rate limits - ideal for building and getting to production. Enterprise adds production rate limits, scheduling priority, volume discounts and dedicated support against a monthly commitment. You can start on Developer and move up when your usage grows.

Why isn't model X available yet, and how fast can you add it?

Every model is optimized to run efficiently on the dataflow hardware rather than just loaded as-is. A new checkpoint of an architecture we already run can go live in hours to days; a new version of a supported family, in days to weeks; a genuinely new architecture takes weeks of engineering. Tell us the model you need and we'll give you a straight answer on timing - and if it's a fine-tune of something we already support, it's fast.

Can I control costs and avoid surprise bills?

Yes. Pricing is pay-as-you-go against published per-token list prices, so cost is predictable and easy to model. You only pay for the tokens you use - no minimums on the Developer tier, no commitments. And you can check your consumption live at any time in the console at cloud.infercom.ai.

Do you offer an uptime SLA on the API?

The shared API is best-effort availability without a contractual SLA - it's designed for building, integrating, and scaling to production. If you need a guaranteed SLA, reserved throughput, and no contention from other users, that's exactly what Dedicated Capacity provides.

What happens under heavy load - can my requests get deprioritized?

On shared infrastructure, higher tiers get scheduling priority, so under peak load a lower-tier request can wait behind enterprise traffic. Requests queue rather than fail in normal operation, and the output speed of a running request is unaffected; only the time to first token grows while it waits. Under extreme overload the queue can reject requests. For capacity no other customer can touch, use Dedicated Capacity.

Is prompt caching available?

Yes, on MiniMax M2.7. Repeated context is served from a cache in the chips' memory, so long prompts start faster; nothing is written to disk. Cached tokens are billed at the standard input rate for now.

How does this compare to running open-source models on GPUs myself?

For latency-sensitive work - agents, chat, voice, coding - the dataflow architecture delivers up to 10x the speed, with no cluster to tune. For pure offline batch generation where latency doesn't matter, GPUs can be more cost-effective, and we'll tell you so. Many teams run a hybrid: Infercom for user-facing inference, GPUs for background jobs.

Need more control?

Same platform, same performance - with more isolation when you need it.

Dedicated Capacity

The same API on single-tenant hardware reserved for you, billed monthly - with guaranteed throughput and custom models.

Learn more: Dedicated Capacity

On-Premises

Your datacenter, your hardware - the ultimate isolation, with an air-gapped option for the strictest requirements.

Learn more: On-Premises

Ready to Build the Future of AI in Europe?

Join forward-thinking organizations deploying sovereign AI with world-class performance