Jebadiah v2.1(27B、9B)发布:用 logits 对候选标签打分的决策模型,开放权重
[P] Pecision models that score every allowed label from the logits: Jebadiah v2.1 (27B, 9B), open weights and self-run benchmark results [P]
Jebadiah v2.1 发布 27B 与 9B 两个开放权重模型,将决策视为闭集打分问题,直接从候选标签 logits 取每个选项的概率,无需生成与采样。
I've been building open models that treat a decision as a closed-set scoring problem rather than text generation. The input is structured context plus a typed question with a fixed set of options. The output is a probability for each option, taken from the candidate-label logits, so there's no generation step to parse and no sampling.
v2.1 updates the 27B and the 9B. On the Decision Index 0.3 public suite (my own runs with the unchanged official scorer, submitted to the board, which adds private tests before ranking anything):
- 27B: 57.03 vs 55.11 for v2. Knowledge +2.34, Language +3.89, Retrieval +2.81, Arts +1.31, Tools -2.10
- 9B: 47.20 vs 44.09. Knowledge +3.67, Language +4.34, Retrieval +5.96, Arts +0.01, Tools -0.84
The regressions are concentrated: When2Call fell 6.38 (27B) and 7.44 (9B). I haven't found the cause yet and the cards don't claim statistical significance for any of these deltas.
Contamination check: the benchmark families that overlap public training sources were scanned against the training data with zero hits. It's a text-match scanner, so paraphrased overlap would get through.
Quantization: I measured agreement with the full weights on 260 held-out questions per build (for example 27B Q8_0 259/260, 9B MLX 4-bit 237/260) and publish the per-question records.
Weights (Apache-2.0, Qwen base): https://huggingface.co/frontier-infra/jebadiah-27b-v2-1 and https://huggingface.co/frontier-infra/jebadiah-9b-v2-1
Results: https://huggingface.co/datasets/frontier-infra/jebadiah-v2-1-decision-index-results
Code and evals: https://github.com/getainode/jebadiah
I'd welcome critique of the evaluation setup, especially the tool-use regression.
submitted by /u/WebDevToday
[link] [留言]
来源:r/MachineLearning · reddit.com