Dedicated Capacity

Guaranteed performance, on the models you choose

Single-tenant AI inference hardware, reserved exclusively for you in our data center in Munich and fully operated by us. No other tenants competing for throughput, a predictable monthly cost, and your own mix of models - including ones beyond the public catalog.

Same platform, more control

The same sovereign platform - reserved for you alone

The Fit

When to choose Dedicated Capacity

It pays off once your workload is high, steady, and production-critical. Below that, the shared API is usually the better deal - and we'll say so.

High, steady volume

Sustained production traffic that consistently pushes past what the shared tiers economically cover.

Guaranteed SLAs

Contractual uptime and response-time commitments, with no contention from other tenants.

Custom models

Run fine-tuned checkpoints (BYOC) or multiple models on one rack, switched in milliseconds.

Isolation

Single-tenant hardware for compliance, predictable performance, or peace of mind.

What You Get

A production inference engine, run for you

Your own high-throughput AI inference plant - serving models at production scale, with sovereign, single-tenant infrastructure that Infercom deploys, operates, and keeps running so your team can focus on building, not ops.

Infrastructure

  • Single-tenant hardware reserved for you - one rack or many
  • Fully operated and maintained by Infercom
  • Guaranteed throughput - no noisy neighbours
  • Custom rate limits, or none at all
  • Models at the precision their makers released them in
  • In our data center in Munich, with no persistent storage of your data

Support & SLAs

  • Dedicated account team
  • Priority support with direct escalation
  • Custom SLAs (uptime, response time)
  • Right-sizing from your real workload
  • ISO 27001 certified operations
  • Deployment typically within ~2 weeks

Models

Run the models you need

A dedicated rack can run models beyond the public API catalog - including ones we bring up specifically for you - plus your own fine-tuned checkpoints via Bring Your Own Checkpoint.

  • The full public catalog
  • Models beyond the public catalog, brought up on request
  • BYOC - fine-tunes of supported architectures load immediately
  • Up to ~700B parameters on a single rack

Fine-tunes of supported architectures load immediately; genuinely new architectures are a commercial conversation. We won't promise a specific custom model without checking it can run.

Available on your rack

  • The full Infercom catalogEvery model you can call on the API
  • Plus models on requestBeyond the public catalog
  • Your own checkpointsFine-tunes of supported architectures (BYOC)
  • Precision as releasedNo quantization today

Fine-tuning itself runs on GPUs - you train, then hand us the weights to serve at full speed.

View full model catalog

One rack, many models

Switch checkpoints in milliseconds

Each chip has three memory tiers: SRAM on the chip for the live computation, HBM in the same package for the weights of the running model, and a large DDR tier for further models and checkpoints. Because the chip addresses its DDR directly, a dedicated rack keeps several models and fine-tuned checkpoints ready and switches between them in milliseconds, versus the seconds to minutes a GPU cluster needs.

That means one rack can serve a whole agentic pipeline - a planner model, an executor, and a specialised fine-tune - switching per step without standing up separate infrastructure for each.

How many models fit on a rack at once depends on their size - we'll size it to your mix.

Checkpoint switch time

Typical GPU reloadseconds-minutes
Dedicated Capacitymilliseconds

Comparison

Dedicated Capacity vs Inference API

Inference APIDedicated Capacity
InfrastructureShared (multi-tenant)Reserved (single-tenant)
BillingPay-per-tokenFlat monthly
Rate limitsTier-basedCustom or none
ContentionShared queue under loadNone - it's all yours
Custom models (BYOC)Catalog onlyCatalog + your checkpoints
SupportTier-basedDedicated team
SLAsBest-effortCustom, guaranteed

See all three options side by side

FAQ

Questions, answered straight

Is dedicated performance the same as the shared API?

Yes - it's the same hardware and the same speeds. What changes is that the capacity is reserved for you alone: no other tenants, no shared queue, no contention. You get consistent throughput because nobody else is using your rack.

When does Dedicated Capacity actually make sense versus pay-per-token?

Two things drive the decision, and only one is about cost. On economics, it makes sense above sustained high utilization - once your steady volume consistently pushes past what the shared tiers cover (the exact crossover depends on your model and traffic pattern); below that, pay-as-you-go on the Inference API is usually the better deal. The other driver has nothing to do with load: if you want your own mix of models, or your own fine-tuned checkpoints (BYOC), running on hardware reserved for you alone, Dedicated is how you get that - at any volume. It's also the answer for guaranteed capacity, SLAs, or when you're past the shared-tier rate limits. If your workload is variable or exploratory and you don't need that control, we'll tell you to stay on consumption.

Can I run my own fine-tuned model?

Yes, if it's a fine-tune of an architecture we already support - it loads and serves immediately. You fine-tune on GPUs and hand us the weights; we serve them on your rack. A genuinely new architecture needs engineering work, which is a commercial conversation. Tell us which model you have in mind and we'll confirm it can run.

Can you fine-tune the model for us?

No - fine-tuning and training run on GPUs, not on this architecture, which is built for inference. You (or your ML partner) fine-tune on GPUs, and we host and serve the resulting checkpoint at full speed. We're transparent about this so there's no mismatch in expectations.

How fast can you switch between models or checkpoints?

Milliseconds, versus seconds to minutes on GPUs. Waiting checkpoints sit in the chip's large DDR tier, which the chip addresses directly, so one rack can hold several fine-tuned checkpoints and switch between them per request, even running different models at different steps of an agentic flow.

Which models can I run, and how large can they be?

Any model in the catalog, plus supported custom checkpoints. On one rack, models up to ~700B parameters run well - a 670B-class model fits comfortably, and MiniMax M2.7 Ultraspeed is a sweet spot. Models above ~1T parameters don't fit on a single rack today.

How many racks do I need?

We size it from your workload - total users, peak concurrency, workload type (interactive, agentic, batch), and throughput targets - worked back from measured concurrency. We'd rather right-size it than oversell you capacity you won't fill.

Can I test before committing?

Yes, and we recommend it. Prove the value on the shared Inference API first (it's technically identical), then we can arrange a time-limited test load before you sign. A dedicated rack is a real commitment - it should follow proven value, not precede it.

How long does deployment take?

A dedicated rack is typically ready in about two weeks. (On-premises hardware in your own datacenter takes longer - see the On-Premises page.)

How is Dedicated Capacity different from On-Premises?

Location and ownership. With Dedicated Capacity, Infercom owns the hardware in our data center in Munich and reserves it exclusively for you - we operate it, you consume it. With On-Premises, you own the hardware in your own datacenter and operate it yourself. Performance is identical; the difference is who holds the keys.

What throughput will I actually get - not the benchmark number?

Real throughput depends on your context length and input/output ratio. Headline tokens-per-second figures are peak numbers on favourable workloads; long-context and output-heavy traffic run lower. We quote throughput against your actual workload, not a single best-case benchmark - that's the honest way to size a rack.

How much does it cost?

Dedicated Capacity is a flat monthly cost per reserved rack, with discounts for longer commitments. Because the right configuration depends on your models, throughput, and contract length, pricing is set per engagement rather than published - talk to us and we'll put real numbers on the table quickly.

Do I have to commit for a long term?

No - longer commitments earn better rates, but we know the model landscape moves fast and you don't want to be locked into last year's setup. We offer shorter terms too. Tell us your risk tolerance and we'll structure something that fits.

Can I resell it or run my own customers' workloads on it?

Yes. It's your reserved capacity - like renting your own server - so you're free to build whatever commercial model you like on top of it, including serving your own clients. System integrators and agencies do exactly this. Talk to us and we'll set it up to fit how you go to market.

Need full ownership?

Consider On-Premises deployment

If a requirement means the hardware must live in your own building - air-gapped environments, or data that cannot leave your network - On-Premises is the ultimate level of isolation.

Learn about On-Premises

Ready to size a dedicated rack?

Tell us your workload and we'll put real numbers on the table quickly.