Moonworks Lunara 提出 Diffusion Mixture Transformer,用不到 10B 激活参数建模艺术智能
Moonworks Lunara: Modeling Artistic Intelligence [R]
Moonworks Lunara 提出 Diffusion Mixture Transformer 架构,激活参数不到 10B,并配套 CAT 训练算法,通过定向采样、图像精修和选择性纳入人类艺术作品迭代更新训练分布。
Lunara introduces a novel Diffusion Mixture Transformer architecture with fewer than 10B active parameters for modeling artistic intelligence in image generation.
Its CAT training algorithm iteratively updates the training distribution through targeted sample acquisition, image refinement, and selective inclusion of human-created artwork inspired by principles of active learning. Semantic variations modify composition while preserving shared content, providing controlled neighborhoods of related training examples.
Evaluation uses 1,000 shared prompts and 8,000 generated images, measuring aesthetic quality, emotional resonance, and content integrity. The seven baselines are GPT-Image-1 Mini, Qwen-Image, AuraFlow, SD 3.5 Turbo, HiDream-I1 Fast, FLUX-Klein-4B, and Z-Image-Turbo.
Under GPT-5.6 Sol evaluation, Lunara leads aesthetic quality at 8.473, versus 8.457 for GPT-Image-1 Mini and 8.366 for Qwen-Image; GPT-Image-1 Mini leads emotional resonance and content integrity.
In the blinded human evaluation, six evaluators assess anonymized image pairs; Lunara achieves the highest mean scores across all three dimensions.
Paper: https://arxiv.org/abs/2609.22272
Evaluation dataset: https://huggingface.co/datasets/moonworks/lunara-art-eval
This release follows the first two open-source dataset releases that reached frontpage of Hugging Face. We hope the findings in this paper can motivate more research in active learning and mixture based architecture for image generation.
submitted by /u/paper-crow
[link] [留言]
来源:r/MachineLearning · reddit.com