Skip to content

Datapoint AI open-sources 2M+ human image preference dataset with 30-model benchmark and $1M grant

Datapoint AI announced the largest open-source human image-preference dataset (2M+ pairwise annotations by real people), a public leaderboard ranking 30 SOTA text-to-image models over 10 use-case categories (marketing, product design, anime, and others), and a $1M data-grant program for academics and non-profits to collect human-preference data for image, video, audio, and other multimodal research. The dataset ships on Hugging Face; the benchmark leaderboard is at trydatapoint.com/benchmark/leaderboard. This is an outside-the-lab alternative preference substrate to the tournament-labeled ERIA-1K / distilled-reasoning-trace pipelines that back current VLM-as-judge reward models.

  • 2M+ human annotations released as an open dataset — datapointai/text-2-image-human-preferences-2m on Hugging Face [tweet OP, thread 2/3].
  • 30 SOTA image generators ranked overall and per-category across 10 use-case categories (marketing, product design, anime, etc.) [tweet OP, thread 2/3].
  • $1M in Datapoint API credits offered as a research grant for academics and non-profits to run their own human-preference collection for image, video, audio, and other multimodal research [thread 3/3, trydatapoint.com/grants].

Announcement post; no methodology paper linked yet. The three released artifacts — the 2M-annotation dataset, the leaderboard, and the grant program — cover a preference-annotation stack from data collection through benchmark scoring. Category-wise rankings on the leaderboard imply a segmentation of the annotation prompt set by use case rather than a single global preference distribution — the same “one-size-fits-all is a hazard” line taken by ERNIE-Image-Aes on the bias-modes side.

  • 2M+ pairwise human annotations (scale claim, no per-category counts disclosed in the announcement).
  • 30 image generators evaluated (specific model list to be read off the leaderboard).
  • 10 use-case categories with individual rankings, in addition to the aggregate ranking.
  • $1M in API credits committed to research grants (deadline / cadence not disclosed).

Sits directly next to ERNIE-Image-Aes: Robust Image Aesthetics Scoring with Balanced Category Generalization (ERNIE-Image-Aes: 8B VLM aesthetic scorer trained on Swiss-tournament labels with category-balanced training on the bias-exposed ERIA-1K benchmark) and Unified Personalized Reward Model for Vision Generation (UnifiedReward-Flex: per-prompt context-adaptive VLM reward for GRPO post-training) as a third substrate for image-preference supervision — this one leaning on scale (2M+) and open licensing rather than curated tournaments or distilled VLM reasoning traces. Also complements Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing (Pico-Banana-400K, text-guided editing dataset) as the generation-side open preference-data release at the ~2M-annotation scale. Whether the Datapoint corpus survives the “one-size-fits-all is a hazard” critique depends on how category-balanced the 10-category segmentation is; that’s an open question until dataset stats land.