Kimi K3 开发者完整指南:Together AI 上的接入与调参
Kimi K3: the complete developer guide
Together AI 发布 Kimi K3 开发者指南,介绍这款 2.8 万亿参数、1M 上下文窗口的开源权重模型在其平台上的完整接入方式。API 兼容 OpenAI,支持 low/high/max 三档 reasoning_effort 与关闭思考的开关,流式响应会分别返回 reasoning_content 与正文。
Together AI 官方给出的 K3 接入指南,覆盖思考档位、工具动态加载与缓存计费,可直接迁移到现有 OpenAI 兼容管线。
Kimi K3 是 Moonshot AI 迄今为止能力最强的模型:一个拥有 2.8 万亿参数的模型,也是全球首个进入 3 万亿参数级别的开源模型。它专为前沿智能工作而设计,例如长周期编程、端到端知识工作和深度推理。它也是首个在 GPT 5.6 Sol 和 Claude Fable 5 层级上竞争的开源权重模型,Together AI 正与 Moonshot 团队直接合作来提供该模型的服务。

已发布的最大开源权重模型
Kimi 团队对扩展规模有着坚定的投入,这一点显而易见:在 2025 年 7 月至 2026 年 7 月的十二个月中,有九个月 Kimi 模型都设定了开源模型规模的上限。K3 拥有 2.8 万亿参数,是迄今为止发布的最大开源权重模型。
内部架构
两项架构更新构成了 K3 的骨干,二者都旨在帮助信息更轻松地流经更长的序列并更深入地进入网络:
- Kimi Delta Attention (KDA):一种混合线性注意力机制,为在超长上下文中扩展注意力提供了高效基础。这是首个支持 1M 上下文长度的 Kimi 模型。
- Attention Residuals (AttnRes):跨模型深度选择性地检索表示,而非统一地累积它们。

在此基础上,Moonshot 通过 Stable LatentMoE 框架进一步推进了混合专家稀疏性,高效地激活 896 个专家中的 16 个。在这种稀疏程度下,每个 token 大约激活 2% 的专家,路由和优化成为首要挑战,因此多项辅助技术使 2.8T 规模下的稳定训练成为可能:
- Quantile Balancing:直接从路由器分数分位数推导专家分配,消除了启发式更新和敏感的平衡超参数。
- Per-Head Muon:扩展 Muon 优化器以独立优化注意力头,从而在大规模下实现更具适应性的学习。
- Sigmoid Tanh Unit (SiTU):改进激活控制。
- Gated MLA:改进注意力选择性。

如何在 Together AI 上使用 Kimi K3
该 API 与 OpenAI 兼容。以下代码片段面向 Together AI,并使用官方 Together Python SDK。
python3 -m pip install --upgrade 'together>=2.0.0'
import os
from together import Together
MODEL = "moonshotai/Kimi-K3"
client = Together(
api_key=os.environ["TOGETHER_API_KEY"],
)
completion = client.chat.completions.create(
model=MODEL,
messages=[{"role": "user", "content": "Introduce Kimi K3 in one sentence."}],
max_tokens=130_000,
)
print(completion.choices[0].message.content)
思考强度
K3 可以通过顶层 reasoning_effort 字段进行配置。支持三个级别:low、high 和 max,默认值为 max。在 Together 上,也可以通过标准的 reasoning={"enabled": False} 开关关闭思考。
# Adjust depth: "low" | "high" | "max"
completion = client.chat.completions.create(
model=MODEL,
reasoning_effort="max",
messages=[{"role": "user", "content": "Prove that the square root of 2 is irrational."}],
max_tokens=8192,
)
# Instant mode, no thinking tokens billed at all
fast = client.chat.completions.create(
model=MODEL,
messages=[{"role": "user", "content": "What is the capital of France?"}],
reasoning={"enabled": False},
max_tokens=256,
)
流式传输
流式响应会分别提供 reasoning_content(思考轨迹)和最终答案的 content 增量。
stream = client.chat.completions.create(
model=MODEL,
messages=[{"role": "user", "content": "Explain why the sky is blue."}],
max_tokens=4096,
stream=True,
)
in_answer = False
for chunk in stream:
if not chunk.choices:
continue
delta = chunk.choices[0].delta
thinking = getattr(delta, "reasoning_content", None) or getattr(delta, "reasoning", None)
if thinking:
print(thinking, end="", flush=True)
if delta.content:
if not in_answer:
print("\n--- answer ---")
in_answer = True
print(delta.content, end="", flush=True)
视觉输入
可以提供多张图像作为输入。Moonshot 还发布了一个视觉推理基准,Perception Bench。
import base64
from pathlib import Path
# Option A: pass an image by URL
IMAGE_URL = "https://raw.githubusercontent.com/pytorch/pytorch/main/docs/source/_static/img/pytorch-logo-dark.png"
image_content = {"type": "image_url", "image_url": {"url": IMAGE_URL}}
# Option B: pass a local image as base64 (uncomment to use)
# image_data = base64.b64encode(Path("image.png").read_bytes()).decode()
# image_content = {"type": "image_url",
# "image_url": {"url": f"data:image/png;base64,{image_data}"}}
completion = client.chat.completions.create(
model=MODEL,
max_tokens=2048,
messages=[{
"role": "user",
"content": [
image_content,
{"type": "text", "text": "Describe this image."},
],
}],
)
视觉限制:
- 图像数量没有限制,但整个请求体必须保持在 100 MB 以下。
- 建议上限:图像为 4K(4096x2160)。更高分辨率会消耗处理时间和 token,却不会提升理解效果。
- Token 成本随分辨率增加。
结构化输出
使用 response_format 搭配 json_schema 和 strict: true 来约束最终 message.content。
import json
completion = client.chat.completions.create(
model=MODEL,
max_tokens=4096,
messages=[{"role": "user", "content": "Ada Lovelace was 36 years old."}],
response_format={
"type": "json_schema",
"json_schema": {
"name": "person",
"strict": True,
"schema": {
"type": "object",
"properties": {"name": {"type": "string"}, "age": {"type": "integer"}},
"required": ["name", "age"],
"additionalProperties": False,
},
},
},
)
person = json.loads(completion.choices[0].message.content)
# -> {'name': 'Ada Lovelace', 'age': 36}
当你只需要语法有效的 JSON 时,更宽松的 {"type": "json_object"} 模式在 Together 上也可用。无论哪种方式,都要让 max_tokens 保持充裕:整个思考轨迹会在第一个受 schema 约束的 token 输出之前消耗完,因此过紧的上限会截断 JSON,而不是截断推理。
工具与 tool_choice
K3 保留标准的工具选择约束。标准循环:在 tools 中声明函数;当模型返回 tool_calls 时,将完整的 assistant 消息追加到历史记录中,然后为每个调用追加一条带有匹配 tool_call_id 的 tool 消息,然后再次调用。在第一轮使用 tool_choice="required" 强制至少调用一次工具,之后切换回 "auto"。更改 tool_choice 不会使前缀缓存失效。
import json
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city.",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name, e.g. Paris"},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
},
"required": ["city"],
"additionalProperties": False,
},
},
}]
def get_weather(city, unit="celsius"):
return {"city": city, "temperature": 21, "unit": unit, "conditions": "sunny"}
messages = [{"role": "user", "content": "What's the weather in Paris?"}]
choice_mode = "required" # force a tool call on turn one
for _ in range(5):
response = client.chat.completions.create(
model=MODEL,
messages=messages,
tools=tools,
tool_choice=choice_mode,
max_tokens=8192,
)
choice = response.choices[0]
message = choice.message
# Append the COMPLETE assistant message, thinking trace included.
messages.append(message.model_dump(exclude_none=True))
if choice.finish_reason != "tool_calls" or not message.tool_calls:
print(message.content)
break
for call in message.tool_calls:
try:
args = json.loads(call.function.arguments)
except json.JSONDecodeError:
args = {}
messages.append({
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps(get_weather(**args)),
})
choice_mode = "auto" # hand control back after the forced turn
动态工具加载
你可以将完整的工具定义(完整名称、描述和参数)放在一条带有 tools 字段且没有 content 的 system 消息中。该工具从该消息的位置开始可用。
关键规则:
- 动态声明使用与顶层 tools 字段完全相同的格式。
- 它们按请求生效,且不会由服务器保留,因此请自行将这条消息保留在后续请求历史中。保留它既能保持工具的可用性,也能保留缓存前缀;丢弃它意味着模型无法再调用该工具,且更改后的前缀可能会错过缓存。
- 将动态声明追加到 messages 末尾不会影响缓存前缀;删除或修改较早的声明可能会损害更改点之后的缓存命中。

大型工具目录的推荐模式:
- 对话开始:只声明一个 search_tools 函数(由你的后端实现)以及少量核心工具,并在系统提示中公布可搜索的领域标签。
- 第一轮:设置 tool_choice: "required" 以强制在回答前进行检索。
- 按需注入:根据检索结果,通过 system 消息插入匹配工具的完整定义。
- 直接调用:模型在后续生成中使用已加载的工具。
- 成本权衡:在对话开始前决定 reasoning_effort。
代码示例:
CATALOG = {
"convert_currency": {
"type": "function",
"function": {
"name": "convert_currency",
"description": "Convert an amount from one currency to another.",
"parameters": {
"type": "object",
"properties": {
"amount": {"type": "number"},
"from_currency": {"type": "string"},
"to_currency": {"type": "string"},
},
"required": ["amount", "from_currency", "to_currency"],
"additionalProperties": False,
},
},
},
}
search_tools = {
"type": "function",
"function": {
"name": "search_tools",
"description": "Search the tool catalog. Tags: finance, travel, files.",
"parameters": {
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"],
"additionalProperties": False,
},
},
}
messages = [{"role": "user", "content": "Convert 100 USD to EUR."}]
# 1. Force retrieval before answering.
first = client.chat.completions.create(
model=MODEL, messages=messages, tools=[search_tools],
tool_choice="required", max_tokens=8192,
)
call = first.choices[0].message.tool_calls[0]
messages.append(first.choices[0].message.model_dump(exclude_none=True))
messages.append({"role": "tool", "tool_call_id": call.id,
"content": json.dumps(list(CATALOG))})
# 2. Inject matching definitions at the TAIL. `tools` field, NO content.
messages.append({"role": "system", "tools": [CATALOG["convert_currency"]]})
# 3. The model calls the freshly loaded tool directly.
second = client.chat.completions.create(
model=MODEL, messages=messages, tools=[search_tools],
tool_choice="auto", max_tokens=8192,
)
print(second.choices[0].message.tool_calls)
# -> convert_currency({"amount":100,"from_currency":"USD","to_currency":"EUR"})
1M 上下文与自动缓存
Together 支持完整的 1M 上下文长度,且上下文缓存是自动的。请让你的长前缀(系统提示、知识库、仓库转储)在多次请求之间保持字节稳定,以便后续调用能够命中缓存。Moonshot 建议将固定的大块上下文(知识文档)放在 messages 数组的最开头,位于 system 消息之前,然后在其后追加问题和回复。
def _get(obj, key, default=None):
if obj is None:
return default
return obj.get(key, default) if isinstance(obj, dict) else getattr(obj, key, default)
usage = completion.usage
reasoning_tokens = _get(_get(usage, "completion_tokens_details"), "reasoning_tokens", 0)
cached_tokens = _get(_get(usage, "prompt_tokens_details"), "cached_tokens",
_get(usage, "cached_tokens", 0))
print(f"prompt={usage.prompt_tokens} cached={cached_tokens} "
f"completion={usage.completion_tokens} thinking={reasoning_tokens}")
# -> prompt=86 cached=64 completion=133 thinking=111
采样参数
采样参数是固定的,你应在请求中省略它们。模型是使用这些参数训练的,不支持设置其他替代值:
- temperature = 1.0
- top_p = 0.95
- n = 1
- presence_penalty = 0
- frequency_penalty = 0
保留思考
K3 是在保留思考历史模式下训练的,因此该轨迹是下一轮所依赖的状态。使用以下方法保留上一轮的思考 token,并将它们转发到后续轮次。

SECRET = "48213"
TRACE = "For the session codeword I will use 48213. Committing to 48213 as the codeword."
messages = [
{"role": "user", "content": "Pick a 5-digit codeword for our session and remember it. "
"Reply with exactly: OK"},
# The trace rides along on the assistant turn. No flag needed.
{"role": "assistant", "content": "OK", "reasoning_content": TRACE},
{"role": "user", "content": "What codeword did you pick? Reply with just the number."},
]
completion = client.chat.completions.create(
model=MODEL,
messages=messages,
max_tokens=4000,
chat_template_kwargs={"preserve_thinking": True}
)
print(completion.choices[0].message.content) # -> 48213
去掉 reasoning_content 这一行,同一个调用每次都会用一个不同的新编造数字来回答。在真实代码中,你从不手写轨迹;你重放模型产生的内容,也就是工具循环中的那一行:
# Turn 1 - let K3 think.
first = client.chat.completions.create(
model=MODEL,
messages=[{"role": "user", "content": "Pick a random 5-digit number and commit to it. "
"Do not tell me. Reply with exactly: OK"}],
max_tokens=4000,
)
# Replay the assistant turn WHOLE. model_dump keeps reasoning_content alongside content.
history = [
{"role": "user", "content": "Pick a random 5-digit number and commit to it. "
"Do not tell me. Reply with exactly: OK"},
first.choices[0].message.model_dump(exclude_none=True),
{"role": "user", "content": "What number did you pick? Reply with just the number."},
]
second = client.chat.completions.create(model=MODEL, messages=history, max_tokens=4000)
print(second.choices[0].message.content)
Kimi K3 定价
Kimi K3 按 token 定价,并设有奖励稳定前缀的缓存命中输入档位:
Kimi K3 定价
| 档位 | 每 1M token 价格 |
|---|---|
| 输入(缓存命中) | \$0.30 |
| 输入(缓存未命中) | \$3.00 |
| 输出 | \$15.00 |
上下文窗口:1,048,576 tokens(1M)。思考 token 按输出计费。

需要内化的两个成本方面:
- 缓存是你的杠杆。在编码工作负载中,命中率超过 90% 时,有效输入成本会趋近于 \$0.30 的下限,但前提是你保持前缀稳定。重构较早的消息或工具声明会破坏这一点。
- 推理按输出计费,且可以调节。思考 token 属于输出 token,价格为 \$15/M,且思考无法完全禁用,但 reasoning_effort 现在有三个级别。max 仍是默认值,因此从不设置该字段的流水线会在每次调用时支付最高推理费用,包括那些微不足道的调用。
Kimi K3 基准测试
在整个评估套件中,Kimi K3 取得了前沿水平的成绩。它在多项编码和智能体基准测试(SWE Marathon、BrowseComp、DeepSearchQA、AutomationBench、OmniDocBench)中领先,在其他测试中与最强的专有模型保持竞争力,同时明显优于另一款受测开放模型 GLM-5.2。在少数基准测试中,它落后于 Claude Fable 5 和 GPT 5.6 Sol,这与 Moonshot 自身对该模型的定位一致。
以下所有 Kimi K3 结果均使用 reasoning effort 设置为 max。
Kimi K3 基准测试
Reasoning effort:max
| 基准测试 |
Kimi K3
max |
Claude Fable 5
max,带 fallback |
GPT 5.6 Sol
max |
Claude Opus 4.8
max |
GLM-5.2
max |
|---|---|---|---|---|---|
| 编码 | |||||
| DeepSWE | 67.5 | 70.0 | 73.0 | 59.0 | 46.2 |
| Program Bench | 77.8 | 76.8 | 77.6 | 71.9 | 63.7 |
| Terminal Bench 2.1 | 88.3 | 84.6 | 88.8 | 84.6 | 82.7 |
| FrontierSWE | 81.2 | 86.6 | 71.3 | 66.7 | 67.3 |
| SWE Marathon | 42.0 | 35.0 | 39.0 | 40.0 | 13.0 |
| PostTrain Bench | 36.6 | 41.4 | 34.6 | 34.1 | 34.3 |
| MLS Bench | 48.3 | 49.9 | 46.2 | 42.8 | 40.4 |
| Kimi Code Bench 2.0(内部) | 72.9 | 76.9 | 64.8 | 71.7 | 64.2 |
| 智能体 | |||||
| GDPval-AA v2(Elo) | 1668 | 1760 | 1748 | 1600 | 1514 |
| BrowseComp | 91.2 | 88.0 | 90.4 | 84.3 | N/A |
| DeepSearchQA(F1) | 95.0 | 94.2 | N/A | 93.1 | N/A |
| Toolathlon-Verified | 73.2 | 77.9 | 74.9 | 76.2 | 59.9 |
| MCP Atlas | 84.2 | 84.7 | 83.6 | 83.6 | 82.6 |
| Automation Bench | 30.8 | 29.1 | 29.7 | 27.2 | 12.9 |
| Job Bench | 52.9 | 57.4 | 46.5 | 48.4 | 43.4 |
| AA-Briefcase(Elo) | 1548 | 1583 | 1495 | 1354 | 1260 |
| APEX-Agents | 41.0 | 43.3 | 39.9 | 39.4 | 35.6 |
| Office QA Pro | 63.3 | 69.9* | 63.2* | 63.9* | 41.4 |
| SpreadsheetBench 2 | 34.8 | 34.7* | 32.4* | 31.6* | 28.1 |
| DECK-Bench(内部) | 73.5 | 73.0 | 74.7 | 66.9 | 68.6 |
| 推理与知识 | |||||
| GPQA-Diamond | 93.5 | 92.6 | 94.1 | 91.0 | 91.2 |
| HLE-Full | 43.5 | 53.3 | 44.5 | 49.8* | N/A |
| HLE-Full w/ tools | 56.0 | 63.0 | 58.0 | 57.9* | N/A |
| 视觉 | |||||
| MMMU-Pro | 81.6 | 81.2 | 83.0 | 78.9 | N/A |
| MMMU-Pro w/ python | 83.4 | 86.5 | 84.6 | 82.7 | N/A |
| CharXiv(RQ) | 84.8 | 88.9 | 84.6 | 80.5 | N/A |
| CharXiv(RQ)w/ python | 91.3 | 93.5 | 89.1 | 89.9 | N/A |
| MathVision | 94.3 | 94.8 | 95.8 | 86.7 | N/A |
| MathVision w/ python | 97.8 | 98.6 | 97.8 | 97.1 | N/A |
| BabyVision w/ python | 85.7 | 90.5 | 88.9 | 81.2 | N/A |
| ZeroBench_main (pass@5) | 23.0 | 23.0 | 17.0 | 17.0 | N/A |
| ZeroBench_main w/ python (pass@5) | 41.0 | 46.0 | 35.0 | 34.0 | N/A |
| WorldVQA ForceAnswer | 51.0 | 56.7 | 41.8 | 39.1 | N/A |
| OmniDocBench | 91.1 | 89.8 | 85.8 | 87.9 | N/A |
| PerceptionBench | 58.5 | 57.2 | 59.7 | 47.2 | N/A |
所有 Kimi K3 结果均使用设置为 max 的推理强度。标有星号(*)的数值是在与基础运行不同的条件下报告的——例如引自外部来源或不同的测试框架。N/A 表示没有已发布的分数。阴影单元格标记该行中的领先结果。各基准测试的精确方法请参见源报告。来源:Kimi K3。
Kimi K3 与前沿模型相比如何
汇总基准测试表格只能说明一部分问题。为了对成本、编码质量和路由行为进行正面比较,我们在 DeepSWE 上让 Kimi K3 与领先的专有模型进行了对比:
常见问题
什么是 Kimi K3? Kimi K3 是 Moonshot AI 的旗舰级 2.8 万亿参数模型,也是首个进入 3 万亿参数级别的开源模型,专为长时程编码、知识工作和推理而构建。
Kimi K3 是开源的吗?
是的。它以开放权重模型的形式发布,Together AI 直接与 Moonshot 团队合作来提供服务。
Kimi K3 的上下文窗口是多少?
100 万 tokens(1,048,576),在 Together AI 上完整支持,并带有自动上下文缓存。
Kimi K3 在 Together AI 上的费用是多少?
每 100 万缓存命中输入 tokens 为 \$0.30,每 100 万缓存未命中输入 tokens 为 \$3.00,每 100 万输出 tokens 为 \$15.00。
可以关闭 Kimi K3 的思考功能吗?
在 Together AI 上,你可以通过 reasoning={"enabled": False} 禁用思考,或通过将 reasoning_effort 设置为 low、high 或 max 来调整推理深度。
Kimi K3 支持视觉吗?
是的。它具备原生视觉能力,每次请求可接受多张图片,只要请求体总大小保持在 100 MB 以下即可。
Kimi K3 已在 Together AI 上线。运行它并部署到生产环境。
让 K3 为你所用,从一次 API 调用开始。
- 运行 Kimi K3 推理:Together AI 上的 Kimi K3 API
- 开始使用 API 构建:阅读文档
来源:Together AI Blog · together.ai