EU HostedNew
MiniMax M2.7 API in Europe
229B frontier reasoning at 428 tokens per second. Fully EU sovereign.
Output Throughput, p50, 10K input / 1K output, single request. Last measured: July 2026. See full benchmarks and methodology
The latest 229B parameter frontier model from MiniMax, running on Infercom's fully EU sovereign infrastructure in Germany. No US hyperscalers. No CLOUD Act exposure. Native multi-agent support, 30% coding improvement over the previous generation.
Get StartedNew in M2.7
M2.7 adds built-in self-critique for better first-attempt results, plus native multi-agent orchestration as a core capability.
Built-in Self-Critique
New in M2.7: The model automatically reviews and refines its outputs before responding. Better first-attempt results on complex coding and reasoning tasks.
Native Agent Teams
New in M2.7: Multi-agent collaboration as a core model capability - not just prompting. Stable role boundaries, adversarial reasoning, and behavioral differentiation built in.
Improved: Software Engineering
56.22% on SWE-Pro (matching GPT-5.3-Codex), 55.6% on VIBE-Pro for full project delivery, 57.0% on Terminal Bench 2 for complex system understanding.
Improved: Professional Skills
97% skill compliance across 40+ complex professional skills. Enhanced Excel, PowerPoint, and Word editing with multi-round revision support.
Measured on Infercom EU Infrastructure
Output Throughput
Time to First Token
Context Window
192K tokens
Server-side p50, 10K input / 1K output, single request
Up to 444 tok/s on shorter prompts. Last measured: July 2026.
Frontier Benchmark Performance
M2.7 achieves top-tier scores across coding, engineering, and ML benchmarks - matching or exceeding GPT-5.3-Codex on SWE-Pro.
56.22%
SWE-Pro
Matches GPT-5.3-Codex
76.5%
SWE Multilingual
Real-world engineering
55.6%
VIBE-Pro
Near Opus 4.6 level
57.0%
Terminal Bench 2
Deep system understanding
66.6%
MLE Bench Lite
Medal rate (9 gold, 5 silver)
52.7%
Multi SWE Bench
Complex codebases
Built for Agentic Workflows
M2.7 excels at long-horizon agent tasks that require autonomous decision-making, tool use, and multi-step reasoning across complex professional domains.
-
Financial Workflows
End-to-end research, Excel modeling, report generation
-
SRE & DevOps
Log analysis, incident response, production debugging
-
Document Processing
Word, Excel, PowerPoint with multi-round editing
-
ML Competitions
66.6% medal rate on MLE Bench Lite (9 gold, 5 silver)
Agentic Coding
56.22% on SWE-Pro
Matches GPT-5.3-Codex on the most demanding software engineering benchmark. Works with Aider, OpenCode, Cline, Cursor, Continue, Goose, Windsurf, and Claude Code.
Multi-Agent Teams
Native Agent Orchestration
Internalized multi-agent collaboration as a native capability. Stable identity across roles, enhanced emotional intelligence, and robust causal reasoning for production decisions.
Pricing
Run a 229B parameter frontier model with transparent, usage-based pricing.
| Model | Input (per 1M) | Output (per 1M) | Context |
|---|---|---|---|
| MiniMax M2.7 Ultraspeed | €0.60 | €2.40 | 192K |
Prices in EUR excl. VAT. EU sovereign deployment with full GDPR compliance.
Why EU Hosting Matters
Running MiniMax through Infercom keeps your requests under EU jurisdiction:
- Your data is processed exclusively in Germany - never leaves EU jurisdiction
- No US CLOUD Act exposure - no American hyperscaler involvement
- Full GDPR compliance with EU-based Data Processing Agreement
- ISO 27001 certified infrastructure owned and operated by Infercom
- No persistent storage - we never train on your data
- ISO 27001 Certified
- GDPR Compliant
- German Datacenter
- Infercom-Owned Hardware
Start Building in Minutes
OpenAI-compatible API. Drop-in replacement for your existing code. Use model name MiniMax M2.7 Ultraspeed in your API calls.
Pay only for the tokens you use.
from openai import OpenAI
client = OpenAI(
api_key="your-infercom-key",
base_url="https://api.infercom.ai/v1"
)
response = client.chat.completions.create(
model="MiniMax-M2.7",
messages=[{"role": "user", "content": "Your prompt here"}],
max_tokens=4096
)
print(response.choices[0].message.content)MiniMax M2.7: Frequently Asked Questions
How fast is MiniMax M2.7 on Infercom?
428 tokens per second output throughput and 690 ms time to first token, both server-side p50 measured at 10K input / 1K output on our production API in Munich. Shorter prompts peak at up to 444 tok/s. For a 229B frontier reasoning model that is up to 10x faster than GPU-based alternatives, and every figure is reproducible with our open-source benchmark tool.
What is the context window of MiniMax M2.7?
192K tokens (196,608), shared between your prompt and the model's response. That is designed for long agentic runs, multi-file code changes and document sets that would otherwise need to be chunked across several requests.
Is MiniMax M2.7 open-weight?
Yes. MiniMax publishes the M2.7 weights openly on Hugging Face. You can put the model into production through our API straight away, and because the weights are open you are never locked into a single provider.
Can I run MiniMax M2.7 in the EU?
Yes. Infercom serves M2.7 from Infercom-owned hardware in a Tier III+ datacenter in Munich, Germany. Requests are processed inside EU jurisdiction with no US CLOUD Act exposure, the infrastructure is ISO 27001 certified, and an EU data processing agreement is available on request. We do not train on your data.
What does MiniMax M2.7 cost?
Pay-per-token with no minimums and no monthly commitment: EUR 0.60 per million input tokens and EUR 2.40 per million output tokens (excl. VAT). You are billed only for what you use, and live usage and spend are visible in the cloud portal.
Learn More
- API Setup Guide Base URL, model IDs, and code examples
- Performance Benchmarks See how MiniMax M2.7 Ultraspeed performs on our EU infrastructure
- Agentic Coding Guide Set up Aider, OpenCode, Cursor with MiniMax
- gpt-oss-120b When raw throughput matters more than frontier reasoning
- EU Sovereign AI Where your data lives and who can reach it
- Pricing Transparent per-token pricing