word 文档在写作和编辑阶段确实方便,但面对电子阅读器或移动设备时,它的阅读体验往往不如专用格式。epub 作为电子书的主流标准,支持自适应排版、字体缩放和目录导航,更适合在 kindle、apple books 等平台上阅读。如果手头有整理好的 word 文档,将其转换为 epub 是一种自然的扩展需求。
本文介绍用 python 完成这一转换的具体做法。
转换工具的选择
python 生态中处理 word 文档的库不少,但能直接输出 epub 格式的选择相对有限。python-docx 专注于 docx 文件的读写,不涉及电子书格式转换;调用 libreoffice 的命令行方式则依赖外部软件安装,在服务器或容器环境中部署不够轻便。
一种更直接的方案是使用独立的文档处理库,通过纯代码完成加载与导出,不需要 microsoft word 或 libreoffice 的参与。
下面以 spire.doc for python 为例说明操作流程。
基础转换
安装方式:
pip install spire.doc
最简单的转换代码只需三步:
from spire.doc import *
from spire.doc.common import *
# 创建文档对象
doc = document()
# 加载 word 文件
doc.loadfromfile("input.docx")
# 保存为 epub
doc.savetofile("output.epub", fileformat.epub)
# 释放资源
doc.close()
fileformat.epub 参数指定了输出格式。转换过程中,word 文档里的标题样式会自动映射为 epub 的章节结构,读者在电子书阅读器中可以通过目录直接跳转。
进阶:添加封面图片
电子书没有封面会显得不够完整,尤其在书库列表中,封面是读者最先看到的内容。该库提供了单独的 savetoepub 方法,可以在保存时指定封面图片:
doc = document()
doc.loadfromfile("input.docx")
# 创建图片对象并加载封面文件
picture = docpicture(doc)
picture.loadimage("cover.png")
# 保存时传入封面参数
doc.savetoepub("output_with_cover.epub", picture)
doc.close()
封面图片支持 png、jpeg 等常见格式。电子书平台通常建议封面宽度不低于 1600 像素,以保证在高分辨率屏幕上的清晰度。
需要注意的几点
免费版本的限制:该库的免费版在读取或写入 word 文件时,限制为 500 个段落和 25 个表格,超出限额的内容会被自动截断。对于超过这个体量的文档,需要评估是否适用。
格式映射的完整性:word 转 epub 并非所有元素都能完美对应。复杂表格、文本框、特定字体效果可能在 epub 中呈现不同。epub 本质上是 xhtml 的打包格式,其排版能力与 word 的页面模型有差异,转换后建议在目标阅读器中实际验证一次。
批量处理的写法:
import os
input_dir = "word_files/"
output_dir = "epub_files/"
os.makedirs(output_dir, exist_ok=true)
for filename in os.listdir(input_dir):
if filename.endswith(".docx") or filename.endswith(".doc"):
doc = document()
doc.loadfromfile(os.path.join(input_dir, filename))
epub_name = os.path.splitext(filename)[0] + ".epub"
doc.savetofile(os.path.join(output_dir, epub_name), fileformat.epub)
doc.close()
每次循环结束后调用 close() 释放文档对象,避免批量处理时内存持续增长。
知识扩展
将 word 文档转换为 epub 电子书,python 生态提供了几种不同层次的方案。选择哪种,取决于你对格式保真度、开发成本和是否需要处理复杂排版的权衡。
方案一:使用 pypandoc(推荐,成本最低)
pandoc 是文档转换领域的“瑞士军刀”,pypandoc 是其 python 封装。它支持从 docx 到 epub 的直接转换,且能保留标题层级、加粗/斜体、列表和基础表格。
1. 安装
pip install pypandoc
pandoc 本身需要单独安装。windows 用户可从 pandoc.org 下载安装包,macos 用 brew install pandoc,linux 用 sudo apt install pandoc。
2. 基础转换代码
import pypandoc
def docx_to_epub_pandoc(input_path, output_path, title=none, author=none):
"""
使用 pandoc 将 docx 转换为 epub。
"""
extra_args = [
'--toc', # 生成目录
'--toc-depth=2', # 目录深度到二级标题
'--metadata', f'title={title or "untitled"}',
'--metadata', f'author={author or "unknown"}',
'--standalone',
]
pypandoc.convert_file(
input_path,
'epub',
outputfile=output_path,
extra_args=extra_args
)
print(f'✅ epub 已生成:{output_path}')调用示例:
docx_to_epub_pandoc(
'input.docx',
'output.epub',
title='我的电子书',
author='作者名'
)3. 进阶:自定义 css 与封面
pandoc 允许通过 --css 和 --epub-cover-image 参数精细控制输出样式。
def docx_to_epub_pandoc_advanced(input_path, output_path, css_path=none, cover_path=none):
extra_args = [
'--toc',
'--toc-depth=3',
'--metadata', 'title=advanced book',
'--metadata', 'author=author name',
]
if css_path:
extra_args.extend(['--css', css_path])
if cover_path:
extra_args.extend(['--epub-cover-image', cover_path])
pypandoc.convert_file(
input_path,
'epub',
outputfile=output_path,
extra_args=extra_args
)注意:pandoc 对复杂表格(合并单元格、嵌套表格)和自定义 word 样式的支持有限。如果文档包含大量此类元素,转换后可能需要手动调整。
方案二:使用商业库(格式保真度最高)
aspose.words for python
aspose.words 是文档处理领域的商业标杆,转换质量最高,但需要许可证(免费版有功能限制)。
import aspose.words as aw
def docx_to_epub_aspose(input_path, output_path):
# 加载许可证(如有)
license = aw.license()
# license.set_license("license.lic") # 取消注释并填入许可证路径
doc = aw.document(input_path)
# 按分页符拆分章节
options = aw.saving.htmlsaveoptions()
options.document_split_criteria = aw.saving.documentsplitcriteria.page_break
options.save_format = aw.saveformat.epub
doc.save(output_path, options)
print(f'✅ epub 已生成:{output_path}')特点:保留原始排版、表格、图片位置和字体样式的能力最强,适合对保真度要求极高的生产环境。
spire.doc for python
spire.doc 的 api 更简洁,免费版有页数限制,可申请 30 天临时许可证。
from spire.doc import document, fileformat
def docx_to_epub_spire(input_path, output_path, cover_image=none):
doc = document()
doc.loadfromfile(input_path)
if cover_image:
from spire.doc import docpicture
picture = docpicture(doc)
picture.loadimage(cover_image)
doc.savetoepub(output_path, picture)
else:
doc.savetofile(output_path, fileformat.epub)
doc.close()
print(f'✅ epub 已生成:{output_path}')方案三:纯 python 实现(python-docx + ebooklib)
如果你需要完全控制转换流程,或者预算为零,可以用 python-docx 读取 word 内容,再用 ebooklib 构建 epub。
1. 安装
pip install python-docx ebooklib beautifulsoup4
2. 完整代码
import os
from docx import document
from docx.enum.text import wd_align_paragraph
from ebooklib import epub
from bs4 import beautifulsoup
def docx_paragraph_to_html(paragraph):
"""将 python-docx 的段落转换为 html。"""
html_parts = []
for run in paragraph.runs:
text = run.text
if not text:
continue
# 转义 html 特殊字符
text = text.replace('&', '&amp;').replace('<', '&lt;').replace('>', '&gt;')
if run.bold:
text = f'<strong>{text}</strong>'
if run.italic:
text = f'<em>{text}</em>'
if run.underline:
text = f'<u>{text}</u>'
html_parts.append(text)
content = ''.join(html_parts)
if not content.strip():
return ''
# 根据 word 样式映射为 html 标签
style_name = paragraph.style.name if paragraph.style else ''
if style_name.startswith('heading 1'):
return f'<h1>{content}</h1>'
elif style_name.startswith('heading 2'):
return f'<h2>{content}</h2>'
elif style_name.startswith('heading 3'):
return f'<h3>{content}</h3>'
elif style_name.startswith('list'):
return f'<li>{content}</li>'
else:
return f'<p>{content}</p>'
def docx_to_epub_pure(docx_path, epub_path, book_title=none, author=none):
"""纯 python 实现:python-docx + ebooklib。"""
doc = document(docx_path)
# 1. 将 word 段落转换为 html 片段
html_fragments = []
in_list = false
for para in doc.paragraphs:
frag = docx_paragraph_to_html(para)
if not frag:
continue
# 列表项包裹
if frag.startswith('<li>'):
if not in_list:
html_fragments.append('<ul>')
in_list = true
html_fragments.append(frag)
else:
if in_list:
html_fragments.append('</ul>')
in_list = false
html_fragments.append(frag)
if in_list:
html_fragments.append('</ul>')
full_html = '\n'.join(html_fragments)
# 2. 创建 epub 书籍对象
book = epub.epubbook()
book.set_identifier(f'id-{os.path.basename(docx_path)}')
book.set_title(book_title or os.path.splitext(os.path.basename(docx_path))[0])
book.set_language('zh')
book.add_author(author or 'unknown')
# 3. 创建章节
chapter = epub.epubhtml(
title='正文',
file_name='content.xhtml',
lang='zh'
)
chapter.content = f'<html><body>{full_html}</body></html>'
book.add_item(chapter)
# 4. 设置目录和书脊
book.toc = (chapter,)
book.spine = ['nav', chapter]
# 5. 添加默认导航样式
style = epub.epubitem(
uid='style_default',
file_name='style/default.css',
media_type='text/css',
content='body { font-family: serif; line-height: 1.6; }'
)
book.add_item(style)
chapter.add_item(style)
# 6. 写入文件
epub.write_epub(epub_path, book)
print(f'✅ epub 已生成:{epub_path}')调用示例:
python
docx_to_epub_pure(
'input.docx',
'output_pure.epub',
book_title='纯 python 转换示例',
author='your name'
)3. 纯 python 方案的局限
- 图片丢失:上述代码未处理 word 中的图片。如需包含图片,需要从
docx中提取inlineshape,保存为文件,并在 html 中以<img>标签引用,同时将图片添加到 epub 的 manifest 中。 - 表格丢失:需要额外遍历
doc.tables并生成<table>html。 - 样式丢失:仅能映射标题层级和粗体/斜体,无法保留字体、颜色、行距等复杂格式。
- 超链接丢失:需要手动解析
paragraph._element中的超链接关系。
小结
word 转 epub 的核心代码量很少,关键在于理解转换过程中格式映射的边界——标题层级、段落样式、基础图片通常能保留,但复杂的页面布局和排版效果不一定能完整迁移。对于以文字为主的文档,转换效果通常可以接受;如果文档包含大量精细表格或特殊版式,转换后需要人工检查。选择工具时,优先考虑能否在目标运行环境(服务器、容器)中正常部署,以及免费方案的限制是否在可接受范围内。
到此这篇关于python文档格式转换之word转为epub的示例代码的文章就介绍到这了,更多相关python word转epub内容请搜索代码网以前的文章或继续浏览下面的相关文章希望大家以后多多支持代码网!
发表评论