跳到正文
r/LocalLLaMA· /u/Abe238·· 4 小时前AI 评分44

DecisionTune 1.0 发布:395M 编码器模型,MLX 上每次短决策约 10 ms(Apache-2.0)

DecisionTune 1.0: a 395M encoder that picks from your options offline, about 10 ms per short decision on MLX (Apache-2.0)

AI 导读

DecisionTune 1.0 发布,这是一个 395M 的决策模型(ModernBERT-large 加 4 KB 评分头),输入状态、问题和选项列表后经一次编码器前向返回各选项概率,不生成文本。

正文

Disclosure: I made this. Sharing it here because it is fully local and small, and I want feedback from people who run models on their own machines.

What it is: a 395M decision model (ModernBERT-large plus a 4 KB scoring head). You give it a state, a question and a list of options. It does one encoder pass and returns a probability for each option, or P(yes) for a yes/no question. It does not generate text.

Why it might be useful in a local stack: the small decisions an agent makes all day (which tool to call, which queue gets a ticket, does this reply answer the question) do not need a large generative model. This handles them on your own machine with no network trip.

Local numbers (our hardware, yours can differ):

  • M5 Pro Mac, MLX backend: median 9.6 ms for a short decision, 1.7 GB of GPU memory.
  • CPU only: about 65 ms per short decision, up to 4.5 GB of memory.
  • Over the full Decision Index run on our laptop: median 25.9 ms, p95 407.8 ms.
  • Weights: 1.58 GB in fp32. Context limit 8,192 tokens. It refuses longer input; it does not truncate.

Backends: PyTorch (default), MLX on Apple silicon (pip install "decision-tune[mlx]", Python 3.11 or newer, selected automatically) and ONNX. Torch and MLX give the same answer on 99.85% of 2,755 questions. Before each release, PyTorch, ONNX and MLX each match the recorded answer on all 50 parity rows.

Offline: after the first download it needs no internet. The package asks before it downloads and checks every file against a SHA-256 manifest.

Quality: 29.57 on Decision Index 0.2.1 (one complete run; a second seed scored 29.13). Strongest area is Tools & Automation at 46.5, up from 28.1 in our 0.9 Preview.

Limits: it only picks from the options you give it. Vague questions with no criteria give weak results, so describe your options ("Shipping: delivery, lost or damaged packages", not "shipping"). Probabilities are not calibrated. English only. Weak at knowledge, math and taste.

Try it:

``` uvx decision-tune ask "Is the customer asking for a refund?" --state "The order arrived broken. I want my money back." ```

or the browser app: pip install "decision-tune[mlx]" then decisiontune app

There is also an MCP server (decisiontune mcp) if you want your local assistant to hand routing and yes/no checks to it.

Model card: https://huggingface.co/decision-tune/decisiontune-1.0 Code: https://github.com/decision-tune/decision-tune Site: https://decisiontune.com/?utm\_source=reddit&utm\_medium=social&utm\_campaign=launch&utm\_content=localllama

If you test it on your own decisions, I would like to hear where it picks wrong.

submitted by /u/Abe238
[link] [留言]

来源:r/LocalLLaMA · reddit.com