Performance

Inference speed, measured

Every number on this page comes from our production API in Munich, with the method and the date next to it. Run the same test yourself.

713tok/s on gpt-oss-120bMeasurements
388ms to the first token, gpt-oss-120bMeasurements
428tok/s on MiniMax M2.7Measurements
up to10xfaster than GPU-based inferenceSources

Output Throughput, p50, 10K input / 1K output, single request. Last measured: July 2026.

Tokens per Second, Measured

These numbers are from our production API with our servers in the EU. Not vendor marketing - real measurements you can reproduce.

gpt-oss-120b

Our fastest

713tok/s · Output Throughput
388ms · Time to First Token
1.789s · End-to-End Latency

Up to 772 tok/s on shorter prompts

MiniMax M2.7 Ultraspeed

Our reasoning model

428tok/s · Output Throughput
690ms · Time to First Token
3.023s · End-to-End Latency

Up to 444 tok/s on shorter prompts

Gemma 4 31B

Native Multimodal

199tok/s · Output Throughput
1,189ms · Time to First Token
6.206s · End-to-End Latency

Native multimodal understanding for images and text

Server-side metrics (p50). Measured using our open-source benchmark tool. Your client-side results will vary based on network location and conditions. Last measured: July 2026.

Fast LLM inference, by workload

Which number decides your speed depends on what you are building. Same measured runs, cut three ways.

Chat and voice

Time to First Token

388 ms

gpt-oss-120b - p50

Conversational and voice interfaces are judged on how fast the first word appears, not on how fast the rest arrives.

Time to First Token

Agents and coding

Output Throughput

713 tok/s

gpt-oss-120b - p50

Agent loops and coding assistants wait for the whole response, so tokens per second sets the pace. For frontier reasoning across long runs, MiniMax M2.7 Ultraspeed measures 428 tok/s with a 192K context window.

Output Throughput

Batch and RAG

End-to-End Latency

1.789 s

gpt-oss-120b - p50

Pipelines care about the total round trip for a complete response. Gemma 4 31B takes image and text in the same request for document and vision workloads.

End-to-End Latency

Fast is half of it. Every number above was measured on hardware we own in Munich - see what EU sovereign actually means.

What happens under load

A request's output speed holds when more requests run in parallel: the chip serves a small batch at a time, and each answer keeps streaming at close to full speed. What grows is the time to first token, when requests have to wait for a place in the next batch. On the shared API that happens at peak times. If your workload needs planned capacity, Enterprise plans come with guaranteed rate limits, and a dedicated rack serves only your traffic.

Dedicated Capacity

Open Source

Run Your Own Benchmark

Our benchmark tool is fully open source. Run it against our API with your API key, or clone the repository and run it locally. Same code, same methodology, your results.

  1. 01

    Synthetic Performance

    Fixed input/output token counts for controlled comparisons across models

  2. 02

    Real Workload Simulation

    Variable request rates mimicking production traffic patterns

  3. 03

    Custom Dataset

    Upload your own prompts and measure performance on your actual workload

  4. 04

    Interactive Chat

    Per-response metrics in a live chat interface - see TTFT and throughput on every reply

How we measure

Three numbers, measured under the same conditions: server-side p50, 10K input and 1K output tokens, one request at a time. We re-measure every quarter.

Time to First Token (TTFT)

How quickly the model starts responding after your request. Critical for interactive applications and chat interfaces.

Output Throughput

Tokens generated per second after the first token. Determines how fast a complete response is delivered to the user.

End-to-End Latency

Total time from request to complete response. Includes TTFT plus full generation time. The number that matters for batch workloads.

Where these numbers come from

All measurements come from our production infrastructure in Munich.

Location
Munich, Germany
Hardware
SambaNova SN40L
Ownership
Hardware we own
Certification
ISO 27001 Certified

Your requests are processed on racks that belong to Infercom, not on rented cloud capacity.

How the hardware works

Speed and energy

Less data movement makes the chip fast and saves energy: a request that finishes sooner draws power for a shorter time.

10 kW

Per rack, typical

About 10 kW in typical use, 7-14.5 kW in inference.

Air-cooled

No liquid cooling

Fits data centers built for regular servers. The chips need no water for cooling.

Up to 5x

Energy efficiency

SambaNova: 2.5x-5.6x against an NVIDIA H200 with Llama 70B, depending on batch size (2025).

Why the chip saves energy

LLM Inference Speed: Frequently Asked Questions

Who has the fastest LLM inference in Europe?

We do not rank other providers; independent comparisons such as Artificial Analysis do that. What we publish is our own measurement: 713 tok/s on gpt-oss-120b, up to 772 tok/s on shorter prompts, measured on our production API in Munich (July 2026). Run the same test against any provider with our open-source benchmark tool.

Run the benchmark yourself

Which model is fastest, and for what?

gpt-oss-120b leads output throughput at 713 tok/s and also has our lowest time to first token at 388 ms, which makes it the default for chat, voice and high-volume agent loops. MiniMax M2.7 Ultraspeed is our 229B frontier reasoning model at 428 tok/s with a 192K context window, for long agentic runs and hard coding tasks. Gemma 4 31B measures 199 tok/s and adds native image and text input. Chat and voice depend on time to first token; agents and coding depend on throughput.

Compare the models

How is this measured - can I reproduce it?

Yes. Every number is a server-side p50 at 10K input and 1K output tokens, single request, measured against our production API with our open-source benchmark tool. Run it online against our API with your own key, or clone the repository and run it locally against your own prompts. Your client-side results will differ depending on network location and conditions.

View the benchmark tool source

What happens to speed when many requests run at once?

Each request keeps streaming at close to full output speed, because the chip serves a small batch at a time. The time to first token grows when requests wait for the next batch, which happens on the shared API at peak times. For planned capacity: Enterprise rate limits or a dedicated rack.

Dedicated Capacity

Is the speed independently verified?

Artificial Analysis continuously measures SambaNova's cloud, and VentureBeat and TechRadar have reported on its speed; a Stanford study of energy efficiency includes the SN40L. Those sources cover the technology. The per-model numbers on this page are our own measurements on our own infrastructure in Munich, which is why we publish the benchmark tool so you can check them.

See the independent benchmarks

Why is dataflow faster than GPUs?

While writing an answer, a GPU runs a model layer as many separate steps and writes the results in between back to memory. A dataflow chip runs the whole layer as one step and keeps those results on the chip, so it spends its memory bandwidth on the weights and answers sooner.

Dataflow vs. GPU architecture, explained

Ready to Build the Future of AI in Europe?

Join forward-thinking organizations deploying sovereign AI with world-class performance