DeepSeek Harness LLM适配器:接入自定义大模型后端
在 DeepSeek Harness 的能力接缝(Capability Seam)架构下,大模型推理能力被抽象为与具体实现无关的统一服务契约。LLM适配器就是该契约的 Service Provider 实现层。通过编写自定义适配器,可以无缝接入任意大模型后端:DeepSeek系列、Ollama/vLLM本地私有化部署、企业代理网关、第三方云端API。
上层Agent循环、工具调用、会话逻辑完全不需要修改,只需要替换适配器插件,即可切换模型后端。本文基于官方文档讲解适配器基类、StreamChunk流式分片协议、错误规范、中止信号、配置挂载与开发踩坑。
一、LLM适配器架构定位
按照三角色能力拆分:
- Definition:
@deepseek‑ai/dsh‑llm,定义LlmAdapter抽象基类、StreamChunk分片类型、LLM服务接口; - Provider:自定义适配器插件,继承抽象基类,对接真实模型后端;
- Consumer:Agent‑Loop 智能体主循环,调用
ctx.llm服务,完全不感知底层模型提供方。
更换模型后端只替换Provider,Agent、Prompt、工具调用逻辑全部保持不变。
二、最小适配器模板:继承LlmAdapter
自定义适配器需要继承 LlmAdapter,重点实现异步生成器方法 stream(),该方法负责把第三方模型返回数据转换为Harness标准的 StreamChunk 流式分片序列。
import type { Context } from '@deepseek‑ai/cordis';
import Schema from '@deepseek‑ai/schemastery';
import { LlmAdapter, type GenerateOptions, type StreamChunk } from '@deepseek‑ai/dsh‑llm';// 自定义适配器类,继承抽象基类
class MyCustomAdapter extends LlmAdapter {
private readonly apiKey: string;
constructor(apiKey: string) {
super();
this.apiKey = apiKey;
}
// 核心:流式生成器,输出标准StreamChunk
async *stream(options: GenerateOptions): AsyncIterable<StreamChunk> {
// 1. 将options.messages、options.tools转换为后端请求格式
// 2. 发起带abort信号的流式HTTP请求
// 3. 将后端响应逐片转换并yield标准StreamChunk
}
}
// 插件配置Schema,Schemastery做配置校验,secret标记密钥
export interface Config {
apiKey: string;
providers: string[];
}
export const Config: Schema<Config> = Schema.object({
apiKey: Schema.string().required().role('secret').description('模型后端API密钥'),
providers: Schema.array(Schema.string()).required().description('适配器对应的提供方标识路由列表')
});
// 声明依赖llm服务
export const name = 'my‑llm‑adapter';
export const inject = ['llm'];
export function apply(ctx: Context, config: Config) {
const adapter = new MyCustomAdapter(config.apiKey);
// 将适配器注册进llm服务,绑定providers路由标识
ctx.llm.registerAdapter(config.providers, adapter);
}
三、StreamChunk 六阶段流式分片协议(强制规范)
stream() 生成器产出的每一个分片必须严格遵守六阶段协议,保证WebUI渲染、工具调用解析、token统计、结束判定可以正常工作。
import { CallId, type StreamChunk } from '@deepseek‑ai/dsh‑llm';async function* demoStream(): AsyncIterable<StreamChunk> {
// 1.块起始:声明块类型与递增索引index从0开始
yield { type: 'block‑start', index: 0, blockType: 'text' };
// 2.增量分片:实时推送文本delta,用于打字机流式输出
yield { type: 'text‑delta', index: 0, text: '你好,' };
yield { type: 'text‑delta', index: 0, text: '我是Harness' };
// 3.块结束:必须携带完整block对象
yield {
type: 'block‑end',
index: 0,
block: { type: 'text', text: '你好,我是Harness' }
};
// 4.工具调用块示例
yield { type: 'block‑start', index: 1, blockType: 'tool‑call' };
yield {
type: 'tool‑call‑delta',
index: 1,
id: CallId('call‑001'),
name: 'execute_bash',
argumentsDelta: '{"command":"ls -la"}'
};
yield {
type: 'block‑end',
index: 1,
block: {
type: 'tool‑call',
id: CallId('call‑001'),
name: 'execute_bash',
arguments: '{"command":"ls -la"}'
}
};
// 5.token消耗统计,必须在finish之前
yield { type: 'usage', usage: { inputTokens: 100, outputTokens: 60 } };
// 6.finish分片:必须作为整个流的最后一个分片
yield { type: 'finish', reason: { kind: 'stop' } };
}
协议强制约束
- 对称闭合:每一个
block‑start必须存在相同index的block‑end - index严格从0开始递增,不可乱序
tool‑call‑delta的argumentsDelta是JSON字符串增量片段,会在block‑end拼装完整参数- usage分片必须出现在finish之前;finish必须是整条流的最后一个分片
该协议同时支持普通文本输出、思维链reasoning块、多工具并发调用,前端既可以使用delta做实时流式渲染,又可以拿到block‑end里完整结构化块用于会话轨迹持久化回放。
四、错误处理、中止信号与归属请求头
1. LlmError标准异常
网络、配额、模型报错时,适配器需要抛出带固定错误码的 LlmError。Agent‑Loop会根据error code执行重试、熔断、上下文截断等自愈策略。
2. AbortSignal中止信号
options.signal 是AbortSignal对象,用户在WebUI点击停止生成时会触发abort。HTTP请求必须带上该signal,保证可以及时断开网络连接,释放资源。
3. attributionHeaders归属头
发起模型API请求时需要合并 attributionHeaders(),携带SDK版本、会话trace标识,用于云端服务审计、链路追踪、配额统计。
import { attributionHeaders, LlmAdapter, LlmError, type GenerateOptions, type StreamChunk } from '@deepseek‑ai/dsh‑llm';export class HttpExampleAdapter extends LlmAdapter {
constructor(private readonly endpoint: string) {
super();
} async *stream(options: GenerateOptions): AsyncIterable<StreamChunk> {
const res = await fetch(this.endpoint, {
method: 'POST',
headers: {
'content‑type': 'application/json',
...attributionHeaders()
},
body: JSON.stringify({
model: options.model,
messages: options.messages,
tools: options.tools
}),
signal: options.signal
}); if (!res.ok) { throw new LlmError(`API异常 ${res.status}`, 'PROVIDER_HTTP_ERROR'); } // ...解析SSE流,yield各个StreamChunk yield { type: 'finish', reason: { kind: 'stop' } }; }
}
五、cordis.yml配置挂载适配器
注册适配器插件,并且配置agent‑loop,让智能体使用我们自定义的provider路由标识。
# 加载自定义LLM适配器插件 - id: my‑llm name: './src/my‑llm‑adapter.ts' config: apiKey: !!js process.env.MY_LLM_API_KEY providers: - my‑private‑cluster将agent‑loop绑定到上面定义的provider id: agent‑loop
name: '@deepseek‑ai/dsh‑agent‑loop'
config:
agents:
- id: main
provider: my‑private‑cluster
model: 'my‑model‑v1'
六、官方参考实现
开源仓库内置两套生产级适配器可供阅读参考:
packages/llm/llm‑deepseek:OpenAI兼容协议适配器,支持DeepSeek‑V3、R1思维链解析、前缀缓存;packages/llm/llm‑pi‑ai:异构API格式接入参考。
七、常见问题FAQ
Q:为什么要区分 block‑start / delta / block‑end?
A:delta用于前端实时打字机渲染;block‑end提供完整校验后的结构化对象,用于会话持久化、轨迹回放;同时支持思维链、多工具调用多块并发推流。
Q:attributionHeaders()有什么作用?
A:携带SDK版本、会话TraceID,用于云端API审计、配额统计、链路问题排查。
Q:如何实现WebUI点击停止生成断开请求?
A:HTTP请求传入 options.signal,用户中断生成会触发该AbortSignal,底层网络连接被终止。
Q:可以同时注册多个LLM适配器吗?
A:可以,不同适配器使用不同的 providers 路由标识;同一个provider只允许一个适配器注册生效。
Q:抛出普通Error和LlmError的区别?
A:必须抛出LlmError并携带标准code;agent‑loop根据code执行重试、熔断、上下文溢出处理;普通Error会被当做未知系统故障。
八、开发最佳实践
- 严格遵循StreamChunk六阶段协议,index递增,保证block‑start与block‑end成对闭合,finish分片必须放在流末尾。
- 所有HTTP请求务必带上options.signal,支持用户主动中止推理,避免后台挂起大量连接。
- 使用
LlmError抛出异常,使用官方规定错误code,让Agent循环可以做故障自愈策略。 - 请求头务必合并
attributionHeaders(),满足可观测性要求。 - 密钥使用Schemastery的
.role('secret'),优先从环境变量读取,禁止硬编码密钥到源码。 - 适配器作为Service Provider,只做模型后端对接;不要修改消息、不要改写工具逻辑,上层业务交给Consumer(agent‑loop)处理。
本篇小结
LLM适配器是Harness能力接缝架构中典型的Service Provider实现:上层Agent‑Loop只依赖llm服务Definition契约,不关心底层是哪一个模型后端。
开发适配器的核心要点:继承LlmAdapter、实现stream异步生成器、严格输出标准StreamChunk分片、处理AbortSignal中止、抛出标准化LlmError异常、注册到对应provider路由。
掌握适配器开发,就可以将任意私有化、本地、第三方大模型接入Harness智能体运行时,无需修改Agent业务逻辑。
0 条笔记