Is it really OpenAI-compatible - can I point my existing code at it?
Yes. It's a drop-in replacement for the OpenAI API: change the base URL to https://api.infercom.ai/v1, use your Infercom API key, and the rest of your code stays the same. We're also Anthropic-compatible, so the Anthropic SDK works too. See side-by-side examples on the API compatibility page.
How do I keep my data in the EU?
Use the EU Sovereign models. They run on our hardware in our data center in Munich, and your data stays in EU jurisdiction. We also store nothing persistently: we don't store prompts or responses on disk, don't train on your data, and don't serve cached answers. The Global Catalog models are processed outside the EU, so choose EU Sovereign models for GDPR-sensitive workloads. More detail on our Trust Center.
Are the models quantized?
No. Today we serve every model at the precision its maker released it in, without quantizing it.
Can we connect privately, without going over the public internet?
Not today: the API is reached over HTTPS on the public internet. If private connectivity is a requirement, talk to us about your network.
How fast is it, really?
Measured on our production API: 428 tok/s on MiniMax M2.7 and 713 tok/s on gpt-oss-120b (July 2026). Output speed per request holds when many requests run in parallel; the time to first token grows when requests queue at peak times. Every response reports its time to first token and tokens per second, so you can check it on your own workload.
What's the difference between the Developer and Enterprise tiers?
Developer is self-service and pay-as-you-go with standard rate limits - ideal for building and getting to production. Enterprise adds production rate limits, scheduling priority, volume discounts and dedicated support against a monthly commitment. You can start on Developer and move up when your usage grows.
Why isn't model X available yet, and how fast can you add it?
Every model is optimized to run efficiently on the dataflow hardware rather than just loaded as-is. A new checkpoint of an architecture we already run can go live in hours to days; a new version of a supported family, in days to weeks; a genuinely new architecture takes weeks of engineering. Tell us the model you need and we'll give you a straight answer on timing - and if it's a fine-tune of something we already support, it's fast.
Can I control costs and avoid surprise bills?
Yes. Pricing is pay-as-you-go against published per-token list prices, so cost is predictable and easy to model. You only pay for the tokens you use - no minimums on the Developer tier, no commitments. And you can check your consumption live at any time in the console at cloud.infercom.ai.
Do you offer an uptime SLA on the API?
The shared API is best-effort availability without a contractual SLA - it's designed for building, integrating, and scaling to production. If you need a guaranteed SLA, reserved throughput, and no contention from other users, that's exactly what Dedicated Capacity provides.
What happens under heavy load - can my requests get deprioritized?
On shared infrastructure, higher tiers get scheduling priority, so under peak load a lower-tier request can wait behind enterprise traffic. Requests queue rather than fail in normal operation, and the output speed of a running request is unaffected; only the time to first token grows while it waits. Under extreme overload the queue can reject requests. For capacity no other customer can touch, use Dedicated Capacity.
Is prompt caching available?
Yes, on MiniMax M2.7. Repeated context is served from a cache in the chips' memory, so long prompts start faster; nothing is written to disk. Cached tokens are billed at the standard input rate for now.
How does this compare to running open-source models on GPUs myself?
For latency-sensitive work - agents, chat, voice, coding - the dataflow architecture delivers up to 10x the speed, with no cluster to tune. For pure offline batch generation where latency doesn't matter, GPUs can be more cost-effective, and we'll tell you so. Many teams run a hybrid: Infercom for user-facing inference, GPUs for background jobs.