很多团队做 Agent,一上来就堆框架:规划器、记忆、工具链、工作流引擎,最后代码几千行,效果却不如一个干净的“模型 + 工具循环”。
Agent 工程的核心,其实就一句话: 让模型决策,让工具执行,让结果回到模型,再决策。
Hermes 模型系列,正是这条链路里很合适的一块“大脑”。
Hermes 是 Nous Research 推出的开源模型系列,以指令跟随、结构化输出和函数调用能力著称。
尤其是 Hermes 2 Pro / Hermes 3 这类版本,对 <tool_call>、JSON 输出、多轮工具调用做了专门优化。
它不追求“什么都自己答”,而是更愿意在需要时输出:
<tool_call>{"name":"get_weather","arguments":{"city":"北京"}}</tool_call>工具执行完后,再把结果用 <tool_response> 回传,模型继续推理。
这就是 Agent 工程最朴素的闭环。
下面代码假设你已经用 vLLM、Ollama 或 Nous API 部署了 Hermes,并提供 OpenAI 兼容接口。 核心逻辑只有 30 行左右:
import json, requests
API = "http://localhost:8000/v1/chat/completions" # vLLM/Ollama 兼容地址
TOOLS = [{
"name": "get_weather",
"description": "查询城市天气",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}]
def get_weather(city):
return {"city": city, "temp": 26, "condition": "晴"}
SYS = "你可以调用工具。工具定义:\n" + json.dumps(TOOLS, ensure_ascii=False) + \
"\n需要时只输出:<tool_call>{\"name\":\"工具名\",\"arguments\":{...}}</tool_call>"
def run_agent(user_input):
messages = [{"role": "system", "content": SYS},
{"role": "user", "content": user_input}]
for _ in range(5):
r = requests.post(API, json={
"model": "hermes",
"messages": messages,
"temperature": 0
}).json()
reply = r["choices"][0]["message"]["content"]
print("模型:", reply)
if "<tool_call>" not in reply:
return reply
call = json.loads(reply.split("<tool_call>")[1].split("</tool_call>")[0])
result = get_weather(**call["arguments"])
messages += [
{"role": "assistant", "content": reply},
{"role": "user", "content": f"<tool_response>{json.dumps(result, ensure_ascii=False)}</tool_response>"}
]
return "达到最大轮次"
print(run_agent("北京今天天气怎么样?"))运行后,你会看到类似流程:
<tool_call> 调用 get_weather;<tool_response> 回传;这就是一个最小可用的 Agent。
上面代码能跑,但不能直接上生产。Agent 工程化要补的是:
Hermes 负责“想”,工程负责“稳”。 代码可以少,但边界、权限、审计不能少。
Hermes 与 Agent 工程的关系,就像大脑与神经系统。 模型提供推理和工具调用能力,工程提供循环、工具、状态和治理。
上面 30 行代码,不是让你直接上线,而是帮你抓住 Agent 的最小闭环。 先跑通它,再逐步叠加记忆、多工具、多 Agent、评估和监控。 Agent 工程的核心,从来不是框架多复杂,而是这条闭环能不能稳定、可控、可解释地跑下去。
原创声明:本文系作者授权腾讯云开发者社区发表,未经许可,不得转载。
如有侵权,请联系 cloudcommunity@tencent.com 删除。