当前位置: 代码网 > it编程>前端脚本>Python > 使用Python代码实现一个生产级Coding Agent

使用Python代码实现一个生产级Coding Agent

2026年10月09日 • Python •我要评论
导读通用 chatbot 还在陪聊,专业的 coding agent 已经在帮工程师定位 bug、修改源码、运行测试并自动提 pr。很多人迷信动辄数十万行代码的庞大开源框架,认为做 agent 必须引

导读

通用 chatbot 还在陪聊,专业的 coding agent 已经在帮工程师定位 bug、修改源码、运行测试并自动提 pr。
很多人迷信动辄数十万行代码的庞大开源框架,认为做 agent 必须引入一堆复杂的链式抽象;但剥离掉那些花哨的封装,顶级 coding agent(如 claude code、hermes agent)的核心内核惊人地简洁与纯粹。
框架是给写 demo 的人用的,协议、工具与状态机才是架构师的武器。300 行纯 python,不依赖任何第三方重量级框架,彻底带你手搓一个真正能跑在生产环境中的 coding agent。(筒子们说了点实话,不喜勿喷哈)

在过去一两年的企业级实践中,很多团队在尝试自研 coding agent 或 devops agent 时,经常陷入一个误区:一上来就引入 langchain、crewai 或 autogen,定义了一大堆复杂的 agent 类、task 链和冗余的 callback。

结果系统上线后,经常出现以下三个令人抓狂的翻车现场:

灾难现象具体翻车表现架构根因
1. 幻觉与伪修复agent 声称“我已经修复了 bug”,缺乏确定性的命令行工具闭环,
(fake fixes)但根本没执行测试,代码甚至跑不通没有把真实测试结果回填到循环中
2. 暴力全文件覆写只改一行代码,却用全文件覆写,缺乏基于精确 diff 的 patch 工具
(destructive write)导致格式错乱并冲掉他人提交和局部修改状态机
3. 上下文无限膨胀读了几个大文件后,token 瞬间耗尽缺乏 token 预算裁剪与滚动摘要,
(context explosion)首字延迟飙升至 10s,服务直接崩溃原始文件内容无节制堆叠在 prompt

真正可用的 coding agent 绝不是靠堆砌提示词生成的,它依赖的是一套严密的执行状态机(conversation loop)与具备确定性反馈的工具集(deterministic toolset)。

一、顶级 coding agent 的底层本质:三位一体架构

如果我们把 claude code、hermes agent 的底层执行流抽丝剥茧,会发现它们的核心架构都可以收敛为极度清晰的“三位一体”模型:

(规划、意图与决策中枢)
runtime tool engine
(bash / read / patch / rg)
memory & context budget
(token 预算、修剪与状态回写)

  • 工具必须具备确定性(deterministic):文件读取必须带行号和截断保护;文件修改必须走 patch 模式(精准锚定替换),严禁整文件覆写;命令执行必须能捕获真实的退出码(exit code);
  • 执行反馈必须真实回流(grounding):agent 修复完代码后,必须通过命令行运行真实测试(如 pytest / npm test),只有当测试真正输出 pass 且退出码为 0 时,才允许判定任务完成;
  • 上下文必须严格控盘(token budgeting):大文件的读取和长命令输出必须设置 hard cap,防止上下文瞬间被撑爆。

小结:
coding agent 的本质不是聊天,而是一个不断尝试、读取真实系统反馈、自我纠错并直到测试全绿的自动化状态机。

二、300 行纯 python 实现生产级 coding agent

以下是完整的、无任何省略号的、可直接运行的 coding agent 完整源码(基于标准库与通用 openai 规范接口,支持接入任何兼容 openai / deepseek / claude 协议的 api):

"""
coding_agent.py - 300 行纯 python 实现生产级 coding agent
包含:react 循环、安全命令执行、行号读取、精准 patch 补丁、上下文预算裁剪
"""

import os
import sys
import json
import shlex
import subprocess
import urllib.request
import urllib.error
from typing import list, dict, any, optional

# ==================== 1. 核心工具集实现 ====================

def tool_run_bash(command: str, timeout: int = 60) -> str:
    """在当前工作目录下执行 shell 命令,捕获真实输出与退出码"""
    # 安全红线拦截
    dangerous_patterns = ["rm -rf /", ":(){ :|:& };:", "mkfs", "dd if=/dev"]
    if any(p in command for p in dangerous_patterns):
        return json.dumps({"error": "security alert: command blocked by safety guardrails", "exit_code": -1})
    
    try:
        proc = subprocess.run(
            command,
            shell=true,
            capture_output=true,
            text=true,
            encoding="utf-8",        # 强制 utf-8 解码(windows 默认 gbk 会炸中文输出)
            errors="replace",        # 非法字节替换为占位符,永不崩溃
            timeout=timeout,
            cwd=os.getcwd()
        )
        output = proc.stdout if proc.stdout else proc.stderr
        # 截断过长输出防止冲垮 context
        if len(output) > 8000:
            output = output[:4000] + "\n... [output truncated due to length] ...\n" + output[-4000:]
        return json.dumps({"output": output, "exit_code": proc.returncode})
    except subprocess.timeoutexpired:
        return json.dumps({"error": f"execution timed out after {timeout} seconds", "exit_code": 124})
    except exception as e:
        return json.dumps({"error": str(e), "exit_code": -1})


def tool_read_file(path: str, offset: int = 1, limit: int = 300) -> str:
    """带行号与分页读取文本文件,格式为 '行号| 内容'"""
    if not os.path.exists(path):
        return json.dumps({"error": f"file '{path}' not found."})
    try:
        with open(path, "r", encoding="utf-8", errors="replace") as f:
            lines = f.readlines()
        
        total_lines = len(lines)
        start_idx = max(0, offset - 1)
        end_idx = min(total_lines, start_idx + limit)
        
        numbered_lines = [f"{i+1:4d}| {lines[i]}" for i in range(start_idx, end_idx)]
        return json.dumps({
            "total_lines": total_lines,
            "showing_lines": f"{start_idx+1}-{end_idx}",
            "content": "".join(numbered_lines)
        })
    except exception as e:
        return json.dumps({"error": str(e)})


def tool_patch_file(path: str, old_str: str, new_str: str) -> str:
    """精准局部替换文件内容,要求 old_str 具备唯一性,严禁全量覆写破坏代码"""
    if not os.path.exists(path):
        return json.dumps({"error": f"file '{path}' does not exist."})
    try:
        with open(path, "r", encoding="utf-8") as f:
            content = f.read()
        
        count = content.count(old_str)
        if count == 0:
            return json.dumps({"error": "target 'old_str' not found in file. check line numbers and indentation."})
        if count > 1:
            return json.dumps({"error": f"target 'old_str' matched {count} times. please include more surrounding context lines for uniqueness."})
        
        new_content = content.replace(old_str, new_str, 1)
        with open(path, "w", encoding="utf-8") as f:
            f.write(new_content)
        
        return json.dumps({"status": "success", "message": f"successfully patched '{path}'."})
    except exception as e:
        return json.dumps({"error": str(e)})


def tool_search_files(pattern: str, base_dir: str = ".") -> str:
    """按文件名查找或关键词正则查找文件路径"""
    matches = []
    for root, _, files in os.walk(base_dir):
        if any(ignored in root for ignored in [".git", "node_modules", "__pycache__", ".venv"]):
            continue
        for f in files:
            if pattern.lower() in f.lower():
                matches.append(os.path.relpath(os.path.join(root, f), base_dir))
    return json.dumps({"matched_files": matches[:50]})

# ==================== 2. 工具契约定义 ====================

tools_schema = [
    {
        "type": "function",
        "function": {
            "name": "run_bash",
            "description": "execute a shell command (e.g. pytest, git, ls, npm test) in the workspace.",
            "parameters": {
                "type": "object",
                "properties": {"command": {"type": "string", "description": "the shell command to run"}},
                "required": ["command"]
            }
        }
    },
    {
        "type": "function",
        "function": {
            "name": "read_file",
            "description": "read file content with line numbers. use offset/limit for large files.",
            "parameters": {
                "type": "object",
                "properties": {
                    "path": {"type": "string", "description": "relative file path"},
                    "offset": {"type": "integer", "description": "start line (1-indexed)"},
                    "limit": {"type": "integer", "description": "max lines to return"}
                },
                "required": ["path"]
            }
        }
    },
    {
        "type": "function",
        "function": {
            "name": "patch_file",
            "description": "replace a unique exact string in a file with new string. always include context lines.",
            "parameters": {
                "type": "object",
                "properties": {
                    "path": {"type": "string", "description": "file path"},
                    "old_str": {"type": "string", "description": "exact text to find and replace (must be unique)"},
                    "new_str": {"type": "string", "description": "replacement text"}
                },
                "required": ["path", "old_str", "new_str"]
            }
        }
    },
    {
        "type": "function",
        "function": {
            "name": "search_files",
            "description": "find files by name pattern in the codebase.",
            "parameters": {
                "type": "object",
                "properties": {"pattern": {"type": "string", "description": "file name substring to search"}},
                "required": ["pattern"]
            }
        }
    }
]

tool_map = {
    "run_bash": tool_run_bash,
    "read_file": tool_read_file,
    "patch_file": tool_patch_file,
    "search_files": tool_search_files
}

# ==================== 3. llm api 调用客户端 ====================

def call_llm(messages: list[dict[str, any]], api_key: str, base_url: str, model: str) -> dict[str, any]:
    """通过标准 http 请求调用兼容 openai 的大模型接口"""
    url = f"{base_url.rstrip('/')}/chat/completions"
    headers = {
        "content-type": "application/json",
        "authorization": f"bearer {api_key}"
    }
    payload = {
        "model": model,
        "messages": messages,
        "tools": tools_schema,
        "tool_choice": "auto",
        "temperature": 0.1
    }
    req = urllib.request.request(url, data=json.dumps(payload).encode("utf-8"), headers=headers)
    try:
        with urllib.request.urlopen(req, timeout=120) as resp:
            return json.loads(resp.read().decode("utf-8"))
    except urllib.error.httperror as e:
        error_msg = e.read().decode("utf-8")
        raise runtimeerror(f"llm api http error {e.code}: {error_msg}")

# ==================== 4. agent 核心执行循环 ====================

class codingagent:
    def __init__(self, api_key: str, base_url: str = "https://api.openai.com/v1", model: str = "gpt-4o"):
        self.api_key = api_key
        self.base_url = base_url
        self.model = model
        self.system_prompt = (
            "you are an expert autonomous software engineering agent.\n"
            "working directory: " + os.getcwd() + "\n"
            "guidelines:\n"
            "1. inspect before modifying: use 'search_files' and 'read_file' to understand codebase structure.\n"
            "2. precision edits: always use 'patch_file' with sufficient unique context lines. never overwrite files completely.\n"
            "3. ground truth verification: always execute relevant unit tests via 'run_bash' after making changes.\n"
            "4. keep working iteratively until tests pass (exit_code=0). only provide your final answer once verified."
        )

    def run(self, user_goal: str, max_turns: int = 15) -> str:
        messages: list[dict[str, any]] = [
            {"role": "system", "content": self.system_prompt},
            {"role": "user", "content": user_goal}
        ]
        
        print(f"\n🚀 [agent started] goal: {user_goal}")
        print("=" * 70)

        for turn in range(1, max_turns + 1):
            print(f"\n🔄 [turn {turn}/{max_turns}] thinking...")
            
            # 上下文预算控制 (超过 30 轮时剔除早期冗余工具输出)
            if len(messages) > 25:
                messages = [messages[0]] + messages[-20:]

            resp = call_llm(messages, self.api_key, self.base_url, self.model)
            choice = resp["choices"][0]
            message = choice["message"]
            messages.append(message)

            # 1. 纯文本回答 (意味着任务完成或需要用户介入)
            if not message.get("tool_calls"):
                content = message.get("content", "")
                print(f"\n✅ [agent finished]:\n{content}")
                return content

            # 2. 工具调用处理
            for tool_call in message["tool_calls"]:
                tool_name = tool_call["function"]["name"]
                tool_args = json.loads(tool_call["function"]["arguments"])
                call_id = tool_call["id"]
                
                print(f"  🛠️  tool call: {tool_name}({json.dumps(tool_args, ensure_ascii=false)})")
                
                # 执行真实工具
                fn = tool_map.get(tool_name)
                if fn:
                    tool_result = fn(**tool_args)
                else:
                    tool_result = json.dumps({"error": f"tool '{tool_name}' not implemented."})
                
                # 截断单次工具回包预览
                preview = tool_result[:150] + "..." if len(tool_result) > 150 else tool_result
                print(f"  📥  result: {preview}")

                messages.append({
                    "role": "tool",
                    "tool_call_id": call_id,
                    "content": tool_result
                })

        return "task stopped: max iterations reached without completion."

# ==================== 5. cli 启动入口 ====================

if __name__ == "__main__":
    if len(sys.argv) < 2:
        print("usage: python coding_agent.py '<task_goal>'")
        sys.exit(1)
        
    api_key = os.getenv("openai_api_key", "your-api-key")
    base_url = os.getenv("openai_base_url", "https://api.openai.com/v1")
    model = os.getenv("agent_model", "gpt-4o")
    
    agent = codingagent(api_key=api_key, base_url=base_url, model=model)
    agent.run(sys.argv[1])

三、关键设计剖析:为什么这 300 行代码能真正干活?

仔细阅读俺上面这段代码,会发现它在工程上解决了自研 agent 最容易暴毙的几个关键问题:

1. 为什么不用write_file,而是坚持patch_file?

在大模型写代码时,最危险的操作就是“整文件覆写”。如果一个文件有 500 行,模型只需要改第 20 行,全文件覆写会导致:

  • 消耗海量输出 token,速度极慢;
  • 模型可能在第 300 行“自作主张”省略一部分代码,导致源码损坏;
  • 冲掉 git 中其他人的并发修改。 patch_file 强制模型提供包含上下文的 old_str 和 new_str,并在代码中严格校验匹配次数(必须严格等于 1)。如果有多次匹配,直接拒绝并要求模型补充更多上下文,从而在数学上保证了修改的原子性与确定性。

2. 为什么需要命令输出的物理截断保护?

当 agent 运行 pytest 或 npm run build 时,有时会输出长达数万行的依赖报错或堆栈日志。如果不加节制地把全部 stdout 灌入 prompt,一次 tool call 就会把上下文窗口打满,导致后续推理直接崩溃。代码中采用“保留头部 4000 字符 + 保留尾部 4000 字符”的切片策略,既保留了初始命令参数,又保留了最终的 exit code 与核心报错栈,信息密度最高。

3. 上下文滑动窗口的软衰减策略

在第 20 轮交互之后,历史中间过程里的很多“中间调试尝试”已经失去了即时价值。代码中的 messages = [messages[0]] + messages[-20:] 实现了极简的记忆窗口截断——保留全局 system 指令,仅保留最近的 20 轮上下文,从而将 token 开销和首字延迟(ttft)稳定在一个恒定的可控区间。

四、生产实战演练:让它自动排查并修复一个真实 bug

我们可以用一个真实的 python 项目来检验这 300 行代码的实战威力。

1. 构造一个有 bug 的业务文件 (calculator.py)

def calculate_tax(income: float, rate: float) -> float:
    # 错误的税率计算逻辑:忘记转换为小数且未处理负数
    if income < 0:
        return income * rate
    return income * (rate / 10)

2. 配套的单元测试 (test_calculator.py)

import pytest
from calculator import calculate_tax

def test_positive_tax():
    assert calculate_tax(10000, 0.2) == 2000.0

def test_negative_income():
    with pytest.raises(valueerror):
        calculate_tax(-100, 0.2)

3. 运行 coding agent

执行命令:

python coding_agent.py "运行 pytest 查看测试失败原因,并修复 calculator.py 中的错误,直到所有测试通过"

4. 控制台真实交互轨迹

以下是在中文 windows 11 + python 3.12 + glm-5.2-cl(openai 兼容协议)环境下的真实运行日志,任务输入与上节命令逐字一致:

🚀 [agent started] goal: 运行 pytest 查看测试失败原因,并修复 calculator.py 中的错误,直到所有测试通过
======================================================================

🔄 [turn 1/15] thinking...
  🛠️  tool call: search_files({"pattern": "calculator"})
  📥  result: {"matched_files": ["calculator.py", "test_calculator.py"]}
  🛠️  tool call: search_files({"pattern": "test"})
  📥  result: {"matched_files": ["test_calculator.py"]}

🔄 [turn 2/15] thinking...
  🛠️  tool call: read_file({"path": "calculator.py"})
  📥  result: {"total_lines": 5, "showing_lines": "1-5", "content": "   1| def calculate_tax(income: float, rate: float) -> float:..."}
  🛠️  tool call: read_file({"path": "test_calculator.py"})
  📥  result: {"total_lines": 9, "showing_lines": "1-9", "content": "   1| import pytest\n   2| from calculator import calculate_tax..."}

🔄 [turn 3/15] thinking...
  🛠️  tool call: run_bash({"command": "cd \"g:\\tmp\\coding-agent-test\" && python -m pytest test_calculator.py -v"})
  📥  result: {"output": "failed test_calculator.py::test_positive_tax - assert 200.0 == 2000.0 ... 2 failed ..."}

🔄 [turn 4/15] thinking...
  🛠️  tool call: patch_file({"new_str": "if income < 0:\n        raise valueerror(\"收入不能为负数\")\n    return income * rate", "old_str": "if income < 0:\n        return income * rate\n    return income * (rate / 10)", "path": "calculator.py"})
  📥  result: {"status": "success", "message": "successfully patched 'calculator.py'."}

🔄 [turn 5/15] thinking...
  🛠️  tool call: run_bash({"command": "cd \"g:\\tmp\\coding-agent-test\" && python -m pytest test_calculator.py -v"})
  📥  result: {"output": "... 2 passed in 0.01s ..."}

🔄 [turn 6/15] thinking...
✅ [agent finished]:
所有测试已通过!
1. 税率计算:将 income * (rate / 10) 改为 income * rate(0.2 表示 20%,无需再除以 10);
2. 负数收入处理:当 income < 0 时抛出 valueerror("收入不能为负数");
修复后两个测试均通过:2 passed in 0.01s

三个值得注意的真实工程细节:

  • 模型没有照剧本走:它先用了两轮 search_files + read_file 主动侦查,第三轮才跑测试——先理解、再验证、后修改,这正是 system prompt 里 "inspect before modifying" 约束真实生效的证据;
  • patch 是全函数级替换:模型提供了完整的新旧函数体做唯一锚点,我们的代码校验匹配次数后一次性替换——严禁全文件覆写的约定被模型严格遵守;
  • 第一次运行时暴露过一个真实 bug:中文 windows 下 subprocess 默认 gbk 解码,命令输出含中文时解码线程直接抛 unicodedecodeerror(这正是本文代码中 encoding="utf-8", errors="replace" 两个参数的由来)——单元测试测不出编码问题,只有真实模型会话踩到中文输出才会炸。

整个过程 6 轮交互,无需人工介入,自动完成了"侦查 → 执行测试 → 捕获报错 → 定位源码 → 打 patch 修复 → 回归测试 → 验证终态"的完整闭环。

五、从 300 行原型到企业级 agent 的演进路线

当然,这 300 行代码是 coding agent 的最小可工作微内核。要将其推向支撑上千研发团队的企业级基础设施,还需要在以下四个维度进行工业化演进:

memory os
跨项目记忆
经验沉淀
  • 接入 memory runtime:挂载我们在本专栏前四篇构建的 memory service,让 agent 能够记住用户的代码规范偏好与历史排错经验;
  • 工具总线协议化(mcp):将单机 python 函数升级为支持跨网络、标准化的 model context protocol,方便对接 github、jira、ci/cd 外部系统;
  • 多 agent 并行派发(subagents):引入后台子代理(batch fan-out),主 agent 负责规划拆解,并行派发多个子 agent 同时在独立沙箱中进行模块编写;
  • 安全沙箱隔离(security sandbox):将命令执行由宿主机 subprocess 替换为 docker 或 gvisor 轻量级隔离沙箱,杜绝越权破坏风险。

总结

  • 抛弃复杂的框架迷信:现代 agent 的本质就是带有确定性反馈的 conversation loop 与状态机;
  • 确定性工具链是成败核心:严禁全量覆写(必须用 patch),严禁只说不做(必须运行真实测试);
  • 代码越纯粹,架构越可控:用最精炼的 300 行代码理解核心,才是迈向企业级全栈 agent 架构师的最佳起点。

以上就是使用python代码实现一个生产级coding agent的详细内容,更多关于python实现生产级coding agent的资料请关注代码网其它相关文章!

赞 (0)

相关文章:

版权声明:本文内容由互联网用户贡献,该文观点仅代表作者本人。本站仅提供信息存储服务,不拥有所有权,不承担相关法律责任。 如发现本站有涉嫌抄袭侵权/违法违规的内容, 请发送邮件至 2386932994@qq.com 举报,一经查实将立刻删除。

发表评论

验证码:
Copyright © 2017-2026  代码网 保留所有权利. 粤ICP备2024248653号
站长QQ:2386932994 | 联系邮箱:2386932994@qq.com