跳到正文
Hugging Face Blog·· 2 小时前同新闻AI 评分63

Liquid AI 开源 d1-3B 与 d1-omni-600M 端侧决策模型

Multimodal open d1 decision models for the edge

AI 导读

Liquid AI 发布 d1 决策模型家族的两款开源权重模型 d1-3B 和 d1-omni-600M(实验性)。d1-3B 在 Decision Index 0.2.1 上得分 48.57,为 10B 以下最佳,支持文本与图像;d1-omni-600M 支持文本加图像或文本加音频,在七个公开数据集上均分 78.4,超过参数量四倍的 Decider 2B。

同一新闻,精选展示《Liquid AI 发布 d1-3B 与 d1-omni-600M 开放权重决策模型》

正文 · AI 翻译

今天,我们发布 d1 决策模型家族中的两个开放决策模型:d1-3B 和 d1-omni-600M(实验性)。

  • Decision Index 0.2.1 上 10B 以下的最佳决策模型:d1-3B 得分为 48.57,领先于所有 4B 和 9B 模型以及 Decider 35B-A3B(47.11)。
  • 多模态:d1-3B 支持文本和图像,而 d1-omni-600M 支持文本和图像或文本和音频
  • 快速:d1-3B 在 NVIDIA Jetson AGX Thor 上回答一个问题需 16 ms,在 Jetson AGX Orin 上需 26 ms,在 Jetson Orin Nano 上需 50ms

我们如何为边缘构建决策模型

这些开放的 d1 决策模型基于我们的 Liquid Foundation Models(LFMs)构建。与我们的生成模型不同,决策模型不产生 token,而是在单次前向传播中给出答案。

d1-3B 和 d1-omni-600M 从两个非常不同的骨干网络训练而来:

  • d1-3B 从 LFM2.5-VL-3B 训练而来,这是我们最新的 VLM,为仅解码器架构。它接受文本和图像作为输入。
  • d1-omni-600M 从 LFM2.5-Encoder-350M 训练而来,这是一个双向编码器。它增加了视觉和音频编码器以处理全部三种模态。它接受文本和图像,或文本和音频作为输入。该模型目前处于早期研究发布阶段,仍在进一步开发中。

基准测试结果

我们在七个公开数据集上对 d1-3B 和 d1-omni-600M 进行了基准测试,涵盖阅读理解、毒性检测、意图分类、医学问答和跨语言理解。d1-3B 取得平均分 82.9,为表中最高,高于 Decider 4B。d1-omni-600M 得分为 78.4,以仅四分之一的参数量超越了 Decider 2B(77.1)。

基准测试 d1-omni-600M d1-3B Decider 2B Decider 4B
SQuAD 2.0 74.0 83.3 67.7 76.0
Civil Comments 95.8 93.3 93.6 92.8
MASSIVE intent 86.1 86.9 81.1 88.3
PubMedQA 61.3 68.3 65.7 63.3
BoolQ 77.7 86.3 87.3 89.0
XNLI 74.7 85.6 85.0 88.6
PAWS-X 79.5 76.4 59.5 69.8
平均 78.4 82.9 77.1 81.1

我们验证了 d1-3B 在标准视觉基准上保留了其 LFM2.5-VL-3B 骨干网络的视觉能力,并且 d1-omni-600M 能够处理全部三种模态。我们未报告任何视觉或音频基准,因为 Decision Index v0.3 仅包含一个私有视觉划分,而音频决策基准目前仍是一个开放问题。

速度

我们与 NVIDIA 合作,在 NVIDIA GeForce RTX 4090、NVIDIA Jetson AGX Thor、Jetson AGX Orin 64 GB 和 Jetson Orin Nano 上基于 NVIDIA 技术栈对 d1-3B 进行了评估。由于 d1-omni-600M 是早期研究发布版本,我们在本次发布中不报告其任何速度数据。

边缘推理。d1-3B 在每台被测设备上回答单个问题均在 50 ms 以内。三个问题耗时仅为单个问题的 1.3 倍,其中 AGX Thor 从 16 ms 增至 20 ms。

一个问题 3 个问题 3.4K-token 状态 384px 图像 64 个状态,打包
Apple M5 Pro 30 ms 41 ms 640 ms 62 ms 78 / s
Jetson AGX Thor 16 ms 20 ms 220 ms 35 ms 262 / s
Jetson AGX Orin 64 GB 26 ms 35 ms 560 ms 83 ms 110 / s
Jetson Orin Nano 50 ms 73 ms 1,640 ms 202 ms 38 / s

GPU 推理。在 GPU 上,d1-3B 在两个平台上回答一个问题均在 10 ms 以内,处理一张 384px 图像均在 18 ms 以内。

一个问题 3 个问题 3.4K-token 状态 384px 图像 64 个状态,打包
NVIDIA RTX 4090 8 ms 21 ms 102 ms 17 ms 475 / s
AMD MI325X 9 ms 14 ms 44 ms 18 ms 1,106 / s

如何使用开放的 d1 决策模型

当你需要快速、结构化的决策(包括多模态输入)时,请选用 d1 决策模型。d1-3B 在其规模下提供了最高的决策质量,而 d1-omni-600M 则适用于对占用空间有要求的场景。

安装依赖项(需要 transformers>=5.14):

pip install "transformers>=5.14" torch torchvision pillow

这些模型自带代码,因此使用 trust_remote_code=True 加载:

import io
import urllib.request

import torch
from PIL import Image
from transformers import AutoModel

device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
model = AutoModel.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True,
                                  dtype=torch.float32 if device == "cpu" else torch.bfloat16).to(device)

# Several named questions over one text state, answered in one pass
questions = {
    "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
    "team": {"type": "choice", "instructions": "Which team should handle this?",
             "criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
                          "fraud": "Suspected unauthorised use"}},
    "urgency": {"type": "score", "instructions": "How urgent is this?",
                "criteria": ["Can wait", "Today", "Blocking the customer now"]},
}
print(model.system_one("I was charged twice this month, please refund one of them.", questions))

# An image as the whole state
url = "http://images.cocodataset.org/val2017/000000039769.jpg"  # two cats on a sofa
photo = Image.open(io.BytesIO(urllib.request.urlopen(url).read()))
print(model.system_one(None, {"cats": {"type": "choice", "instructions": "How many cats are there?",
                                       "criteria": {"one": "One", "two": "Two", "more": "Three or more"}}},
                       images=[photo]))

# Many requests, packed together with no padding
tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))

为简洁起见,我们仅包含 d1-3B 的示例。有关如何运行 d1-omni-600M 的说明,请参阅 d1-omni-600M 模型卡。

开始使用开放的 d1 决策模型

两个决策模型均为开放权重,现已在 Hugging Face 上提供:

我们迫不及待想看到你构建的作品。

引用

如果你使用了这项工作,请引用发布博客:

@article{liquidAI2026opend1,
  author  = {Liquid AI},
  title   = {Open d1: Edge decision models for text, vision, and audio},
  journal = {Liquid AI Blog},
  year    = {2026},
  note    = {www.liquid.ai/blog/open-d1},
}

来源:Hugging Face Blog · huggingface.co