MiMo-V2-Pro (a.k.a. Hunter Alpha) — Artificial Analysis model card
Artificial Analysis benchmark / model card for MiMo-V2-Pro, the production name of the “Hunter Alpha” stealth model that appeared anonymously on OpenRouter on March 11, 2026. Created by Xiaomi’s MiMo team (led by former DeepSeek researcher Luo Fuli), the model is a 1T-parameter, 1M-context reasoning model marketed as a “full-stack model family built for the Agent era.” On the Artificial Analysis Intelligence Index it scores 49 (vs. a peer-class median of 36), at 3.00 per 1M input/output tokens, with 64 tok/s output speed and a 2.50s TTFT. Proprietary weights as of release; Xiaomi has stated they plan to open-source “when the models are stable enough.”
Key claims
Section titled “Key claims”- MiMo-V2-Pro scores 49 on the Artificial Analysis Intelligence Index, well above the peer-class average of 36 [§Intelligence].
- Pricing is 3.00 / 1M output tokens (peer median 8.00); total cost to evaluate on the AA Intelligence Index was $350.87 [§Pricing, §Intelligence Index Token Use & Cost].
- The model is very verbose: it generated 77M output tokens to complete the AA Intelligence Index, vs. a peer-class median of 35M [§Intelligence Index Token Use & Cost].
- Output speed is 64 tok/s (below the peer median of 70 t/s); TTFT is 2.50s (better than the peer median 2.67s) [§Speed, §Latency].
- The model is a reasoning model using extended thinking / chain-of-thought, text-only input/output, with a 1M-token context window [§Model summary, §Context Window].
- The weights are proprietary as of the card; Xiaomi has not disclosed total or active parameter counts on the AA page [§Model Size FAQ].
Method
Section titled “Method”The page is a benchmark / pricing card, not a research artifact — there is no methodology disclosed. The relevant background, confirmed independently after the card was published: “Hunter Alpha” and “Healer Alpha” appeared as anonymous “stealth models” on OpenRouter on 2026-03-11, with a 1T-parameter / 1M-context spec sheet and no developer attribution. On 2026-03-19 Luo Fuli (head of Xiaomi MiMo, former DeepSeek researcher) publicly identified Hunter Alpha as “an early internal test build of MiMo-V2-Pro,” and announced the broader MiMo-V2-Pro + Omni + TTS family as Xiaomi’s first full-stack model family targeted at agentic use. The AA card here is the post-reveal benchmark profile for the production MiMo-V2-Pro reasoning variant, with Xiaomi’s own API as the serving provider.
The AA Intelligence Index aggregates the following sub-evaluations (each linked from the card): GDPval-AA (agentic real-world tasks), Terminal-Bench Hard (agentic coding/terminal), τ²-Bench Telecom (agentic tool use), AA-LCR (long-context reasoning), AA-Omniscience accuracy + non-hallucination rate, Humanity’s Last Exam, GPQA Diamond, SciCode, IFBench, CritPt (physics), APEX-Agents-AA (long-horizon agentic), MMMU-Pro (visual reasoning). The card surfaces an aggregated index score but the per-eval table is not reproduced in the body fetched.
Results
Section titled “Results”- AA Intelligence Index: 49 (peer-class median 36; class = reasoning models in a similar price tier) [§Intelligence].
- Cost to run the Intelligence Index: $350.87, generating 77M output tokens (peer median 35M) — i.e. roughly 2× the verbosity of peer reasoning models on the same benchmark suite [§Intelligence Index Token Use & Cost].
- Output speed: 63.8 tok/s on Xiaomi’s API (peer median 69.7) [§Speed].
- TTFT: 2.50s (peer median 2.67s) [§Latency].
- Context window: 1.0M tokens (text only; no image input) [§Context Window, §FAQ].
- Pricing: 3.00 output / 1.20/1M); “very competitive” vs. peer reasoning-tier medians [§Pricing, §FAQ].
Why it’s interesting
Section titled “Why it’s interesting”Tracking the “stealth-launch → benchmark profile → reveal” cadence is itself useful: Hunter Alpha is the second prominent anonymous-on-OpenRouter launch this year, and the pattern (anonymous benchmark, then attribution after positioning) is now a reproducible playbook. The model is the latest entrant in the closed-but-API-accessible reasoning-model cohort that the wiki has been tracking as a comparison baseline (Open foundation-model releases explicitly enumerates this cohort under Open Questions). It also adds a non-Western frontier-class reasoning model alongside the Qwen / GLM / DeepSeek / Kimi line — Xiaomi entering the agentic-reasoning frontier from a base of smartphones and EVs is a new vector. Of independent interest: the 77M-token AA evaluation footprint at 2× peer verbosity is a quantitative datapoint for what “extended thinking” actually costs at the new generation of reasoning models, complementing the verbosity / token-budget analyses tracked in Reasoning RL and Inference-Time Scaling.
See also
Section titled “See also”- Open foundation-model releases — MiMo-V2-Pro sits in the “closed-but-API-accessible” cohort enumerated under that page’s Open Questions; Xiaomi has stated intent to open-source “when stable.”
- Reasoning RL — MiMo-V2-Pro is a chain-of-thought reasoning model; the AA card reports verbosity 2× peer-median, a relevant datapoint for trace-length analyses.
- Inference-Time Scaling — Total AA-Intelligence-Index cost of $350.87 and 77M output tokens is a concrete inference-compute datapoint for the cohort.
- Gemini Embedding 2: SOTA multimodal embedding model (Google product announcement) — Gemini Embedding 2 announcement (another closed-but-API-accessible release from the same month).
- Kimi K2.5: Visual Agentic Intelligence — Kimi K2.5 ships a comparable “Agent era” full-stack family from Moonshot; useful peer for product framing.