Skip to content

True Story of Pangu — Huawei Noah's Ark whistleblower statement

Anonymous statement posted to GitHub on 2025-07-06 by an author claiming to be an employee of Huawei’s Noah’s Ark Lab on the Pangu LLM team. The author alleges that several models in the Pangu series shipped under “from-scratch on Ascend NPUs” framing were in fact continued-pretraining of competitor weights — most concretely that Pangu-MoE 72B (Pangu Pro MoE) was initialized from Qwen 2.5 14B and trained to scrub the parameter-distribution fingerprint, and that an earlier Pangu 135B V2 was a continued-pretraining of Qwen 1.5 110B with added layers / widened FFN to pad parameter count. The statement explicitly defends Pangu Ultra (Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs, 135B V3) and a 718B MoE as real from-scratch Ascend runs done by the “fourth-column” core team rather than the alleged-laundering “small-model lab.” Surfaces a benchmark-integrity question for the broader Pangu series: how much of the published Pangu evaluation table is real, and which architectural/training contributions (such as depth-scaled init) retain their evidentiary weight after the disclosure.

  • Pangu 135B V2 alleged to be continued-pretraining of Qwen 1.5 110B: author claims V2 was built by adding layers and widening FFN on a Qwen 1.5 110B checkpoint, with a Pangu-π-style mechanism appended to pad the parameter count to ~135B; reports the new V2 had 82 layers (vs the original 135B’s 107) and parameter distributions matching Qwen 1.5 110B nearly exactly, with model class names still containing “Qwen” in the internal codebase [statement §“王云鹤和他的小模型实验室出手了”].
  • Pangu-MoE 72B (Pangu Pro MoE) alleged to be continued-pretraining of Qwen 2.5 14B: author asserts the 72B MoE was initialized from Qwen 2.5 14B and trained over substantial token volume specifically to wash out the parameter-distribution similarity, including deliberately training on noisy data to scrub watermarks; claims that during this washing the team spent more compute on the cover-up than a fresh from-scratch training run would have cost [statement §“小模型实验室也开启了第二次主要的套壳行动”].
  • Pangu Ultra 135B V3 (Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs) and a 718B MoE explicitly defended as real from-scratch runs on Ascend NPUs, attributed to a different internal team (“fourth column / 四纵”) from the alleged-laundering group (“small-model lab / 王云鹤’s lab”); the author affirms the 13.2T-token / no-loss-spike training curve in the V3 tech report as truthful [statement §“135B V3是我们四纵团队当时的骄傲”].
  • Benchmark-evaluation evidence chain: cites HonestAGI’s public analysis comparing Pangu-MoE 72B and Qwen 2.5 14B internal-tensor signatures as triggering the wider controversy; author acknowledges HonestAGI’s analysis was not airtight, leaving room for Huawei’s official rebuttal, but defends the conclusion on the basis of insider observation [statement §“HonestAGI的事情出来后”].
  • Token sequence laundering details: claims that to wash the Qwen 1.5 110B tokenizer signature out of V2, the team used a stitched tokenizer built from two existing vocabularies and continued-trained ≥1T tokens before the embedding distribution matched the new tokenizer — the alleged mechanism is reported to have had a “more serious bug” than known publicly [statement §“71B和135B模型都有一个巨大的硬伤”].
  • Internal authorship and review process critique: author states they will request removal from the author list of Pangu technical reports they signed, and alleges the internal review process at Huawei did not catch (or chose not to catch) the alleged provenance issues despite cross-team awareness [statement §“我也在申请从盘古部分技术报告的作者名单中移除”].
  • Self-identification disclaimer: author concedes they cannot publish internal records as evidence because doing so would expose them via Huawei’s data-leak tracing, and asks readers to weigh the listed insider-only details (team names, project codenames, internal product timelines, Suzhou off-site logistics) as identity proof rather than relying on documentary evidence [statement §“我以生命,人格和荣誉发誓”].

The artifact is a long-form Chinese-language whistleblower statement posted as a single README in a fresh GitHub repository, with subsequent appendices added as the author received external attention and pressure. The statement is structured as: (a) self-authentication via insider-only details (team head names, internal codenames like “盘古智子”, Suzhou off-site mechanics, project structure under the “四野 / four-armies” organizational name); (b) chronological account of the Pangu model series across three generations (V1: 71B/135B from-scratch on 910A with poor tokenizer; V2: alleged-laundered from Qwen 1.5 110B; V3: real from-scratch on 910B by a different team); (c) parallel account of the alleged MoE laundering pipeline (224B MoE attempts, then Pangu-MoE 72B from Qwen 2.5 14B, then alleged 718B continued from DeepSeek-V3); (d) commentary on Huawei’s internal process and the author’s personal decision to leave the company.

The artifact is not a technical analysis paper — there are no comparable-tensor figures, no replicated benchmark numbers, no signature-matching plots. It is a sworn first-person testimony whose evidentiary mode is inside-knowledge specificity rather than reproducible measurement. The strongest external-evidence link is to HonestAGI’s public tensor-similarity analysis between Pangu-MoE and Qwen 2.5, which the author cites but does not repeat.

There are no quantitative results in the conventional sense. The artifact’s “outputs” are:

  • A claim graph (which Pangu models are alleged to be laundered, from which Qwen checkpoint, with what cover-up mechanism).
  • A defense graph (Pangu Ultra 135B V3 and a 718B MoE explicitly carved out as real from-scratch work).
  • A process-and-culture critique of Huawei’s internal incentive structure that the author argues enabled the alleged behavior.
  • A list of insider-only details offered as identity authentication.

The downstream measurable impact (to Luma) is on citation trust — whether the contributions of papers in the Pangu series (particularly Pangu Ultra’s DSSN + depth-scaled init, and Pangu Pro MoE’s reported training recipe) should be weighted normally, weighted down pending replication, or treated as suspect by default until independently verified.

For the wiki, this is a rare integrity-event artifact: a whistleblower disclosure that asks a research team to re-weight citations to a specific paper series rather than re-think a technical claim. Luma’s published interest in sandwich-norm and depth-scaled init means the team needs to decide which Pangu-derived ideas survive this. Jiaming’s intake note captures the right calibration: sandwich norm itself is overdetermined by independent prior work and the architectural-stability cluster on Training stability at scale (where multiple papers — A Unified View of Attention and Residual Sinks: Outlier-Driven Rescaling is Essential for Transformer Training, How to Set the Learning Rate for Large-Scale Pre-training?, mHC: Manifold-Constrained Hyper-Connections — propose related primitives independently of Pangu Ultra), so adopting it from the Pangu Ultra paper is fine; depth-scaled init rests load-bearing on Pangu Ultra’s reported no-loss-spike training curve at 135B / 13.2T and so needs independent replication before being adopted as a recommendation. The artifact also sharpens an open question for Open foundation-model releases: how to read benchmark tables on first-party releases when whistleblower allegations exist on the same model series, and what the minimum-viable-release-package should include to make laundering claims falsifiable (tensor-signature plots? tokenizer-evolution training-time tracking? a release-time provenance attestation?).