Perceptron Egocentric — first embodied-reasoning offering, SOTA over Gemini 3.5 Flash and Gemini Robotics-ER 1.6 annotation pipelines
Perceptron AI (Bellevue-based physical-AI lab, founded Nov 2024 by ex-Meta researchers Armen Aghajanyan and Akshat Shrivastava) launched Perceptron Egocentric on 2026-07-09 — the first in a stated series of “Embodied Reasoning offerings,” pitched as beating the best robotic-annotation pipelines built on Gemini 3.5 Flash and Gemini Robotics-ER 1.6. Access is partner-gated (early-access, closed model). No paper, no report, no numbers were published; only the launch tweet plus a 57-second demo video. The July 22 follow-up from CEO Armen Aghajanyan (“Our embodied reasoning models are getting even better”) is a hype pointer that adds no technical content. This page exists to place a stake in the wiki for the Egocentric line before more concrete artifacts land.
Key claims
Section titled “Key claims”- Perceptron Egocentric is a VLM specialized for egocentric / first-person video annotation for robot learning, positioned as an annotation-pipeline replacement rather than an end-to-end VLA policy [tweet body].
- The stated headline is that it “achieves SOTA over the best robotic annotation pipelines built on Gemini 3.5 Flash and Gemini Robotics-ER 1.6” — a comparison against composed pipelines on top of those Google models, not against the base models in isolation [tweet body].
- Access is partner-only at launch; no benchmark numbers, no evaluation protocol, no cost per token, and no model card are disclosed in the launch tweet or the follow-up [tweet body; @ArmenAgha follow-up].
Method
Section titled “Method”Not disclosed in the tweet or the follow-up. Public context (not part of this artifact): Perceptron’s earlier flagship Mk1 (announced 2026-05-12) is a closed-source video-and-embodied-reasoning model with structured spatial primitives (point / box / polygon / track / clip) as first-class outputs alongside text, priced at 1.50 per million input/output tokens — designed as the perception layer feeding downstream VLA policies (subtask boundaries for hierarchical VLAs, success/failure labels for reward models, action-conditioned annotations for world model training, quality scores for episode filtering). Egocentric is framed by the launch tweet as the first of a series of specialized embodied-reasoning offerings on top of that lineage.
Results
Section titled “Results”None disclosed. The tweet’s SOTA claim is asserted without a table, a benchmark name, or a target task. The demo video is 57 seconds and not analyzable from the text-only capture.
Why it’s interesting
Section titled “Why it’s interesting”- Slots directly into the VLM-as-annotator lane of VLM-as-Evaluator. The productized-verifier / productized-annotator shape currently has one filed commercial entry — Instance Labs — Verifying Robot Learning Episode Success (per-episode success verdicts). Perceptron Egocentric is a second commercial entry aimed one layer upstream: not verifying rollouts but labeling human/robot egocentric footage for downstream VLA training. Same open question applies (no disclosed calibration against human raters), same partner-only access pattern.
- Directly competes with the Google stack that several filed VLA data pipelines already lean on. Segmenting Robot Video into Actionable Subtasks (WGO-Bench) uses Gemini-3.5-Flash on contact sheets at $2.64/hr for subtask segmentation; the WGO-Bench pipeline is exactly the kind of “annotation pipeline built on Gemini 3.5 Flash” the Perceptron tweet claims to beat. Also complements the deployment-time perception layer that industrial VLAs need — sibling in framing to 3D-Object Perception Transformer (3PT) (3PT for 6-DoF pose) but pushed into the annotation-time half of the stack.
- Counter-position to end-to-end VLA scaling. The recipe board on VLA Models currently tracks seven levers (action pretraining at scale, clean teleop, unified VLM pointing, frozen WFM + action expert, native video+action pretraining, memory-into-weights via TTT, and scale-of-embodiment-free-UMI-data). Perceptron’s public position — that a specialized VLM annotator on top of a raw video corpus can replace the frontier-model annotation pipelines feeding those recipes — is an eighth upstream lever: change the substrate that all seven recipes consume, not the recipe itself.
See also
Section titled “See also”- VLM-as-Evaluator — Perceptron Egocentric is a productized VLM annotator, same lane as Instance Labs and Macrodata’s WGO-Bench pipeline.
- Instance Labs — Verifying Robot Learning Episode Success — the only other filed commercial productization of VLM-as-evaluator for robot learning; Instance verifies rollouts, Perceptron labels raw video.
- Segmenting Robot Video into Actionable Subtasks (WGO-Bench) — the Gemini-3.5-Flash-based subtask segmentation pipeline that Perceptron’s SOTA claim implicitly targets.
- VLA Models — Perceptron Egocentric sits upstream of every recipe on the VLA board, in the data-substrate lane.
- Synthetic Training Data — annotation pipelines for VLA training data.