Skip to content

Perceptron Egocentric — first embodied-reasoning offering, SOTA over Gemini 3.5 Flash and Gemini Robotics-ER 1.6 annotation pipelines

Perceptron AI (Bellevue-based physical-AI lab, founded Nov 2024 by ex-Meta researchers Armen Aghajanyan and Akshat Shrivastava) launched Perceptron Egocentric on 2026-07-09 — the first in a stated series of “Embodied Reasoning offerings,” pitched as beating the best robotic-annotation pipelines built on Gemini 3.5 Flash and Gemini Robotics-ER 1.6. Access is partner-gated (early-access, closed model). No paper, no report, no numbers were published; only the launch tweet plus a 57-second demo video. The July 22 follow-up from CEO Armen Aghajanyan (“Our embodied reasoning models are getting even better”) is a hype pointer that adds no technical content. This page exists to place a stake in the wiki for the Egocentric line before more concrete artifacts land.

  • Perceptron Egocentric is a VLM specialized for egocentric / first-person video annotation for robot learning, positioned as an annotation-pipeline replacement rather than an end-to-end VLA policy [tweet body].
  • The stated headline is that it “achieves SOTA over the best robotic annotation pipelines built on Gemini 3.5 Flash and Gemini Robotics-ER 1.6” — a comparison against composed pipelines on top of those Google models, not against the base models in isolation [tweet body].
  • Access is partner-only at launch; no benchmark numbers, no evaluation protocol, no cost per token, and no model card are disclosed in the launch tweet or the follow-up [tweet body; @ArmenAgha follow-up].

Not disclosed in the tweet or the follow-up. Public context (not part of this artifact): Perceptron’s earlier flagship Mk1 (announced 2026-05-12) is a closed-source video-and-embodied-reasoning model with structured spatial primitives (point / box / polygon / track / clip) as first-class outputs alongside text, priced at 0.15/0.15 / 1.50 per million input/output tokens — designed as the perception layer feeding downstream VLA policies (subtask boundaries for hierarchical VLAs, success/failure labels for reward models, action-conditioned annotations for world model training, quality scores for episode filtering). Egocentric is framed by the launch tweet as the first of a series of specialized embodied-reasoning offerings on top of that lineage.

None disclosed. The tweet’s SOTA claim is asserted without a table, a benchmark name, or a target task. The demo video is 57 seconds and not analyzable from the text-only capture.

  • Slots directly into the VLM-as-annotator lane of VLM-as-Evaluator. The productized-verifier / productized-annotator shape currently has one filed commercial entry — Instance Labs — Verifying Robot Learning Episode Success (per-episode success verdicts). Perceptron Egocentric is a second commercial entry aimed one layer upstream: not verifying rollouts but labeling human/robot egocentric footage for downstream VLA training. Same open question applies (no disclosed calibration against human raters), same partner-only access pattern.
  • Directly competes with the Google stack that several filed VLA data pipelines already lean on. Segmenting Robot Video into Actionable Subtasks (WGO-Bench) uses Gemini-3.5-Flash on contact sheets at $2.64/hr for subtask segmentation; the WGO-Bench pipeline is exactly the kind of “annotation pipeline built on Gemini 3.5 Flash” the Perceptron tweet claims to beat. Also complements the deployment-time perception layer that industrial VLAs need — sibling in framing to 3D-Object Perception Transformer (3PT) (3PT for 6-DoF pose) but pushed into the annotation-time half of the stack.
  • Counter-position to end-to-end VLA scaling. The recipe board on VLA Models currently tracks seven levers (action pretraining at scale, clean teleop, unified VLM pointing, frozen WFM + action expert, native video+action pretraining, memory-into-weights via TTT, and scale-of-embodiment-free-UMI-data). Perceptron’s public position — that a specialized VLM annotator on top of a raw video corpus can replace the frontier-model annotation pipelines feeding those recipes — is an eighth upstream lever: change the substrate that all seven recipes consume, not the recipe itself.