Skip to content

MiMo-V2-Pro (a.k.a. Hunter Alpha) — Artificial Analysis model card

Artificial Analysis benchmark / model card for MiMo-V2-Pro, the production name of the “Hunter Alpha” stealth model that appeared anonymously on OpenRouter on March 11, 2026. Created by Xiaomi’s MiMo team (led by former DeepSeek researcher Luo Fuli), the model is a 1T-parameter, 1M-context reasoning model marketed as a “full-stack model family built for the Agent era.” On the Artificial Analysis Intelligence Index it scores 49 (vs. a peer-class median of 36), at 1.00/1.00 / 3.00 per 1M input/output tokens, with 64 tok/s output speed and a 2.50s TTFT. Proprietary weights as of release; Xiaomi has stated they plan to open-source “when the models are stable enough.”

  • MiMo-V2-Pro scores 49 on the Artificial Analysis Intelligence Index, well above the peer-class average of 36 [§Intelligence].
  • Pricing is 1.00/1Minputtokensand1.00 / 1M input tokens** and **3.00 / 1M output tokens (peer median 1.55/1.55 / 8.00); total cost to evaluate on the AA Intelligence Index was $350.87 [§Pricing, §Intelligence Index Token Use & Cost].
  • The model is very verbose: it generated 77M output tokens to complete the AA Intelligence Index, vs. a peer-class median of 35M [§Intelligence Index Token Use & Cost].
  • Output speed is 64 tok/s (below the peer median of 70 t/s); TTFT is 2.50s (better than the peer median 2.67s) [§Speed, §Latency].
  • The model is a reasoning model using extended thinking / chain-of-thought, text-only input/output, with a 1M-token context window [§Model summary, §Context Window].
  • The weights are proprietary as of the card; Xiaomi has not disclosed total or active parameter counts on the AA page [§Model Size FAQ].

The page is a benchmark / pricing card, not a research artifact — there is no methodology disclosed. The relevant background, confirmed independently after the card was published: “Hunter Alpha” and “Healer Alpha” appeared as anonymous “stealth models” on OpenRouter on 2026-03-11, with a 1T-parameter / 1M-context spec sheet and no developer attribution. On 2026-03-19 Luo Fuli (head of Xiaomi MiMo, former DeepSeek researcher) publicly identified Hunter Alpha as “an early internal test build of MiMo-V2-Pro,” and announced the broader MiMo-V2-Pro + Omni + TTS family as Xiaomi’s first full-stack model family targeted at agentic use. The AA card here is the post-reveal benchmark profile for the production MiMo-V2-Pro reasoning variant, with Xiaomi’s own API as the serving provider.

The AA Intelligence Index aggregates the following sub-evaluations (each linked from the card): GDPval-AA (agentic real-world tasks), Terminal-Bench Hard (agentic coding/terminal), τ²-Bench Telecom (agentic tool use), AA-LCR (long-context reasoning), AA-Omniscience accuracy + non-hallucination rate, Humanity’s Last Exam, GPQA Diamond, SciCode, IFBench, CritPt (physics), APEX-Agents-AA (long-horizon agentic), MMMU-Pro (visual reasoning). The card surfaces an aggregated index score but the per-eval table is not reproduced in the body fetched.

  • AA Intelligence Index: 49 (peer-class median 36; class = reasoning models in a similar price tier) [§Intelligence].
  • Cost to run the Intelligence Index: $350.87, generating 77M output tokens (peer median 35M) — i.e. roughly 2× the verbosity of peer reasoning models on the same benchmark suite [§Intelligence Index Token Use & Cost].
  • Output speed: 63.8 tok/s on Xiaomi’s API (peer median 69.7) [§Speed].
  • TTFT: 2.50s (peer median 2.67s) [§Latency].
  • Context window: 1.0M tokens (text only; no image input) [§Context Window, §FAQ].
  • Pricing: 1.00input/1.00 input / 3.00 output / 1.00cachehit,allper1Mtokens(blended7:2:11.00 cache-hit, all per 1M tokens (blended 7:2:1 → 1.20/1M); “very competitive” vs. peer reasoning-tier medians [§Pricing, §FAQ].

Tracking the “stealth-launch → benchmark profile → reveal” cadence is itself useful: Hunter Alpha is the second prominent anonymous-on-OpenRouter launch this year, and the pattern (anonymous benchmark, then attribution after positioning) is now a reproducible playbook. The model is the latest entrant in the closed-but-API-accessible reasoning-model cohort that the wiki has been tracking as a comparison baseline (Open foundation-model releases explicitly enumerates this cohort under Open Questions). It also adds a non-Western frontier-class reasoning model alongside the Qwen / GLM / DeepSeek / Kimi line — Xiaomi entering the agentic-reasoning frontier from a base of smartphones and EVs is a new vector. Of independent interest: the 77M-token AA evaluation footprint at 2× peer verbosity is a quantitative datapoint for what “extended thinking” actually costs at the new generation of reasoning models, complementing the verbosity / token-budget analyses tracked in Reasoning RL and Inference-Time Scaling.