Datapoint AI open-sources 2M+ human image preference dataset with 30-model benchmark and $1M grant
Datapoint AI announced the largest open-source human image-preference dataset (2M+ pairwise annotations by real people), a public leaderboard ranking 30 SOTA text-to-image models over 10 use-case categories (marketing, product design, anime, and others), and a $1M data-grant program for academics and non-profits to collect human-preference data for image, video, audio, and other multimodal research. The dataset ships on Hugging Face; the benchmark leaderboard is at trydatapoint.com/benchmark/leaderboard. This is an outside-the-lab alternative preference substrate to the tournament-labeled ERIA-1K / distilled-reasoning-trace pipelines that back current VLM-as-judge reward models.
Key claims
Section titled “Key claims”- 2M+ human annotations released as an open dataset —
datapointai/text-2-image-human-preferences-2mon Hugging Face [tweet OP, thread 2/3]. - 30 SOTA image generators ranked overall and per-category across 10 use-case categories (marketing, product design, anime, etc.) [tweet OP, thread 2/3].
- $1M in Datapoint API credits offered as a research grant for academics and non-profits to run their own human-preference collection for image, video, audio, and other multimodal research [thread 3/3, trydatapoint.com/grants].
Method
Section titled “Method”Announcement post; no methodology paper linked yet. The three released artifacts — the 2M-annotation dataset, the leaderboard, and the grant program — cover a preference-annotation stack from data collection through benchmark scoring. Category-wise rankings on the leaderboard imply a segmentation of the annotation prompt set by use case rather than a single global preference distribution — the same “one-size-fits-all is a hazard” line taken by ERNIE-Image-Aes on the bias-modes side.
Results
Section titled “Results”- 2M+ pairwise human annotations (scale claim, no per-category counts disclosed in the announcement).
- 30 image generators evaluated (specific model list to be read off the leaderboard).
- 10 use-case categories with individual rankings, in addition to the aggregate ranking.
- $1M in API credits committed to research grants (deadline / cadence not disclosed).
Why it’s interesting
Section titled “Why it’s interesting”Sits directly next to ERNIE-Image-Aes: Robust Image Aesthetics Scoring with Balanced Category Generalization (ERNIE-Image-Aes: 8B VLM aesthetic scorer trained on Swiss-tournament labels with category-balanced training on the bias-exposed ERIA-1K benchmark) and Unified Personalized Reward Model for Vision Generation (UnifiedReward-Flex: per-prompt context-adaptive VLM reward for GRPO post-training) as a third substrate for image-preference supervision — this one leaning on scale (2M+) and open licensing rather than curated tournaments or distilled VLM reasoning traces. Also complements Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing (Pico-Banana-400K, text-guided editing dataset) as the generation-side open preference-data release at the ~2M-annotation scale. Whether the Datapoint corpus survives the “one-size-fits-all is a hazard” critique depends on how category-balanced the 10-category segmentation is; that’s an open question until dataset stats land.
See also
Section titled “See also”- VLM-as-Evaluator — human-labeled preference data is the training substrate for VLM-as-judge reward models; this release enters that substrate at 2M+ scale.
- Synthetic Training Data — preference labels over generated images are a canonical form of synthetic-adjacent training data (real human labels over real-model outputs).
- Open foundation-model releases — dataset + benchmark + grant program shipped together as one bundle, matching the multi-artifact release pattern this concept tracks.
- ERNIE-Image-Aes: Robust Image Aesthetics Scoring with Balanced Category Generalization — closest sibling on the image-preference-data axis; ERNIE-Image-Aes uses curated Swiss-tournament labels vs Datapoint’s crowd-scale annotations.
- Unified Personalized Reward Model for Vision Generation — UnifiedReward-Flex’s DPO half is trained on preference pairs; a Datapoint-scale corpus is a plausible drop-in.
- Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing — comparable open large-scale image-side preference/editing dataset at ~400K scale.