DeepSeek VLM enters grayscale testing — Deli Chen quote-RT of Xiaokang Chen
DeepSeek engineer Deli Chen (@victor207755822) on Apr 29, 2026 quote-retweets Xiaokang Chen (@PKUCXK, also DeepSeek), whose post signals that DeepSeek’s multimodal model has entered grayscale (limited rollout) testing. Deli’s framing is “The little whale can now see (in grayscale testing),” signalling a first public-facing test of vision capabilities on the DeepSeek stack. No model name, no benchmarks, no API endpoint, no tech report — only the existence of a VLM in private rollout is disclosed.
Key claims
Section titled “Key claims”- DeepSeek has a multimodal (vision-capable) model in grayscale rollout as of Apr 29, 2026 [tweet body].
- The announcement is made by DeepSeek engineers Deli Chen and Xiaokang Chen on X, not via a model card, blog post, or arXiv preprint at the time of the tweet [tweet metadata].
- The accompanying image (a screenshot of what appears to be VLM output) is the only public artifact; no model identifier or release channel is named [tweet attached image].
Method
Section titled “Method”Not applicable — the tweet is a teaser/announcement, not a technical artifact. The quote-RT chain is Xiaokang Chen → Deli Chen, both DeepSeek-affiliated; the underlying work is whichever multimodal model DeepSeek’s “multimodal colleagues” have been training. Grayscale testing in this context means limited, gated access for select users before public release.
Results
Section titled “Results”None disclosed in the tweet. The image attached to the post is presented without quantitative claims; no benchmark numbers, no comparison to closed frontier VLMs (GPT-5.x, Gemini 3, Claude 4.x), and no comparison to other open VLMs (Qwen3-VL, Kimi K2.5, HunyuanImage 3.0) are made.
Why it’s interesting
Section titled “Why it’s interesting”DeepSeek up to now has shipped strong language and OCR models — DeepSeek-OCR 2: Visual Causal Flow (DeepSeek-OCR 2) is the most recent narrow-perception release, and DeepSeek-V4 collection release (Flash + Pro, up to 1.6T) (DeepSeek-V4 Flash + Pro up to 1.6T) is the latest LLM family — but no general-purpose VLM is public yet. A first DeepSeek VLM would complete the open-frontier-VLM cohort that already includes Qwen3-VL (Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking), Kimi K2.5, HunyuanImage 3.0 (HunyuanImage 3.0 Technical Report), and Molmo2 (Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding). The “grayscale testing” framing also suggests DeepSeek may follow a closed-rollout-first → open-weights-later pattern, which would be a departure from their usual day-0 open-weights releases (DeepSeek-V3, R1, OCR 2). Worth tracking for whether the eventual release lands as a paper, weights, or hosted-only product.
See also
Section titled “See also”- DeepSeek-OCR 2: Visual Causal Flow — prior DeepSeek perception release; this VLM likely shares the same visual encoder lineage
- DeepSeek-V4 collection release (Flash + Pro, up to 1.6T) — DeepSeek-V4 LLM family; the LLM backbone candidate for this VLM
- Open foundation-model releases — concept page tracking the open-VLM cohort this would join
- Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking — Qwen3-VL series, the closest open-VLM peer
- HunyuanImage 3.0 Technical Report — HunyuanImage 3.0, current largest open multimodal model