Skip to content

DeepSeek VLM enters grayscale testing — Deli Chen quote-RT of Xiaokang Chen

DeepSeek engineer Deli Chen (@victor207755822) on Apr 29, 2026 quote-retweets Xiaokang Chen (@PKUCXK, also DeepSeek), whose post signals that DeepSeek’s multimodal model has entered grayscale (limited rollout) testing. Deli’s framing is “The little whale can now see (in grayscale testing),” signalling a first public-facing test of vision capabilities on the DeepSeek stack. No model name, no benchmarks, no API endpoint, no tech report — only the existence of a VLM in private rollout is disclosed.

  • DeepSeek has a multimodal (vision-capable) model in grayscale rollout as of Apr 29, 2026 [tweet body].
  • The announcement is made by DeepSeek engineers Deli Chen and Xiaokang Chen on X, not via a model card, blog post, or arXiv preprint at the time of the tweet [tweet metadata].
  • The accompanying image (a screenshot of what appears to be VLM output) is the only public artifact; no model identifier or release channel is named [tweet attached image].

Not applicable — the tweet is a teaser/announcement, not a technical artifact. The quote-RT chain is Xiaokang Chen → Deli Chen, both DeepSeek-affiliated; the underlying work is whichever multimodal model DeepSeek’s “multimodal colleagues” have been training. Grayscale testing in this context means limited, gated access for select users before public release.

None disclosed in the tweet. The image attached to the post is presented without quantitative claims; no benchmark numbers, no comparison to closed frontier VLMs (GPT-5.x, Gemini 3, Claude 4.x), and no comparison to other open VLMs (Qwen3-VL, Kimi K2.5, HunyuanImage 3.0) are made.

DeepSeek up to now has shipped strong language and OCR models — DeepSeek-OCR 2: Visual Causal Flow (DeepSeek-OCR 2) is the most recent narrow-perception release, and DeepSeek-V4 collection release (Flash + Pro, up to 1.6T) (DeepSeek-V4 Flash + Pro up to 1.6T) is the latest LLM family — but no general-purpose VLM is public yet. A first DeepSeek VLM would complete the open-frontier-VLM cohort that already includes Qwen3-VL (Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking), Kimi K2.5, HunyuanImage 3.0 (HunyuanImage 3.0 Technical Report), and Molmo2 (Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding). The “grayscale testing” framing also suggests DeepSeek may follow a closed-rollout-first → open-weights-later pattern, which would be a departure from their usual day-0 open-weights releases (DeepSeek-V3, R1, OCR 2). Worth tracking for whether the eventual release lands as a paper, weights, or hosted-only product.