Skip to content

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World-Centric Foundation GUI Agents

Qwen-UI-Agent is Alibaba MAI-UI’s next-generation foundation GUI agent, positioned as one model that operates real phones, computers, and web browsers while also handling Deep Research. The technical report emphasizes that a GUI agent needs to be more than a GUI-specialist — it must preserve the base Qwen3.5-27B’s general reasoning, multimodal understanding, and tool-use strengths. Author-reproduced benchmarks span three axes: general multimodal/knowledge/math (MMMU-Pro 72.4, MathVision 82.8, MMLU-Pro 86.5), agentic/tool-use (Tau2-Bench 89.9, Terminal-Bench 2.0 avg 50.1, BFCL-v4 74.2), and deep-research browsing (BrowseComp 64.1, BrowseComp-ZH 75.0). Qwen-UI-Agent’s headline is that a GUI-focused fine-tune of Qwen3.5-27B can retain — and in several cases lightly exceed — the base model’s general/agentic scores while decisively beating GUI-specialist peers (UI-Venus 30B-A3B, GUI-Owl 32B, OpenCUA-72B) across all three axes.

  • Qwen-UI-Agent is presented as a real-world-centric GUI foundation agent that unifies mobile GUI use (Android apps), desktop computer use, web browsing, and Deep Research retrieval in one model [§Overview].
  • The base model is Qwen3.5-27B; Qwen-UI-Agent scores within noise of Qwen3.5-27B across seven general multimodal / knowledge / math / instruction-following benchmarks (MMMU-Pro 72.4 vs 73.5, RealWorldQA 83.1 vs 83.1, CharXiv-RQ 77.7 vs 76.8, MathVision 82.8 vs 82.0, AI2D 91.1 vs 91.9, MMLU-Pro 86.5 vs 86.0, IFEval strict 90.2 vs 90.4) [Benchmarks Table 1].
  • Against three GUI-specialist baselines (UI-Venus 30B-A3B, GUI-Owl 32B, OpenCUA-72B), all reproduced in-house, Qwen-UI-Agent shows large gaps on the same general benchmarks — e.g. MMMU-Pro 72.4 vs 31.0–39.5 for baselines, MathVision 82.8 vs 26.6–50.6 — supporting the claim that current GUI specialists trade away general reasoning [Benchmarks Table 1].
  • On agentic and tool-use benchmarks, Qwen-UI-Agent leads Qwen3.5-27B by a small margin on most axes: Tau2-Bench 89.9 vs 89.2, Terminal-Bench 2.0 avg-5 50.1 vs 41.1, Claw-Eval avg-3 73.5 vs 66.9 and Pass@3 51.8 vs 41.2, BFCL-v4 74.2 vs 71.3, SkillsBench avg-5 28.0 vs 24.9 [Benchmarks Table 2].
  • On the same agentic axis, GUI specialists collapse: Terminal-Bench 2.0 avg-5 drops from 50.1 (Qwen-UI-Agent) to 0.0–9.0 (UI-Venus / GUI-Owl / OpenCUA); Claw-Eval Pass@3 drops from 51.8 to 0.5–5.5; SkillsBench drops to ~0 [Benchmarks Table 2].
  • For deep-research retrieval, Qwen-UI-Agent scores BrowseComp 64.1 (vs Qwen3.5-27B 61.0) and BrowseComp-ZH 75.0 (vs 62.1), with GUI-specialist baselines not evaluated on these tasks [Benchmarks Table 2].
  • The blog is explicit that all baseline scores are independently reproduced in the authors’ evaluation environment and that harness / judge / simulator / runtime / task-subset settings may differ from official reports; this is flagged with a ”†” mark and a pointer to the full technical report for documented differences [Note under benchmark tables].
  • The showcased end-to-end workflow (“passion-fruit sour-soup beef” recipe research + timed grocery order) demonstrates cross-app, mobile-device execution: the agent first extracts a recipe and its ingredients from Douyin, then carries that information into the Hema app to complete a time-constrained grocery order with a specified delivery window [Case Study].

The blog does not disclose architecture beyond identifying Qwen3.5-27B as the base model, and it does not disclose training data mix, curriculum, or compute. What it does commit to is the evaluation posture: instead of specializing away from general capability (as UI-Venus, GUI-Owl, and OpenCUA appear to) the authors’ recipe preserves Qwen3.5-27B’s general/agentic strengths while adding GUI competence — the benchmark tables are structured to make this specialization-cost visible. The showcased mobile workflow (Douyin → Hema, with a delivery-time constraint) suggests the model is deployed on real Android devices with cross-app state carry rather than a screen-recording sandbox; the report itself is not yet linked from this page, and technical details would need to be pulled from it separately.

  • General multimodal / knowledge / math / IF. MMMU-Pro 72.4, RealWorldQA 83.1, CharXiv-RQ 77.7, MathVision 82.8, AI2D 91.1, MMLU-Pro 86.5, IFEval strict 90.2 — parity with Qwen3.5-27B base within ~1 point on each [Table 1].
  • GUI-specialist gap on general axes. MMMU-Pro delta of +32.9 to +41.4 vs UI-Venus 30B-A3B / GUI-Owl 32B / OpenCUA-72B; similar magnitude gaps on CharXiv-RQ, MathVision, MMLU-Pro [Table 1].
  • Agentic / tool-use. Tau2-Bench 89.9, Terminal-Bench 2.0 avg-5 50.1, Claw-Eval avg-3 73.5 / Pass@3 51.8, BFCL-v4 74.2, SkillsBench avg-5 28.0, QwenClawBench avg-3 44.2 — small lift over Qwen3.5-27B, large lift over GUI specialists [Table 2].
  • Deep research. BrowseComp 64.1 (+3.1 over Qwen3.5-27B), BrowseComp-ZH 75.0 (+12.9 over Qwen3.5-27B); GUI specialists not evaluated on these [Table 2].
  • All numbers are author-reproduced; the report notes that baseline scores in Tables 1–2 were independently evaluated rather than copied from provider reports.

Qwen-UI-Agent extends the tool-use / computer-use design space along an axis no other filed system pushes as hard: specialization without capability collapse. The Terminal-Bench-2.0 result (avg-5 50.1 vs ~0–9 for GUI specialists) is the strongest single datapoint on the page for the claim that current GUI-specialist recipes trade away general agentic capability, and it complements OSGym: Scalable OS Infra for Computer Use Agents (which shows a Qwen-2.5-VL 7B base can be RL’d to competitive OSWorld numbers via full-OS replicas) and VideoAgentTrek: Computer Use Pretraining from Unlabeled Videos (which mines YouTube tutorials for computer-use pretraining trajectories) — those two focus on scaling the training substrate for GUI agents, while Qwen-UI-Agent focuses on not degrading the base. It sits alongside Qwen3.6-Plus: Towards Real World Agents as the mobile/GUI arm of the same “Qwen towards real-world agents” push — Qwen3.6-Plus for desktop coding + hosted API deep-research, this one for mobile / cross-device GUI. The BrowseComp-ZH lift (+12.9 over the base) is the sharpest research-agent datapoint in the Chinese-language pole and complements Tongyi DeepResearch: A New Era of Open-Source AI Researchers and Marco DeepResearch: Unlocking Efficient Deep Research Agents via Verification-Centric Design on that axis. Caveats: no weights, no report link on the landing page, and all numbers are author-reproduced — meaning the specialist-baseline scores in particular should be treated as reproducible-in-principle but not officially validated.