OCR 不可行就用 PyMuPDF 提取文字,支持加密双层 pdf。然后保存到同名的 txt 用 AnyTXT 检索
deepseek有频次限制就换豆包呗。
豆包有限制就换元宝呗。
元宝有限制就换千问呗。
我觉得你最后一个方法就挺好。
不过可以更加批量化一些。
首先批量把所有文件的前四页复制出来。然后写个批量拼接成一页。文件名就和PDF保持一致。
然后挨个图片粘给豆包,让它格式化输出书名。
然后通过自动化提取豆包的输出,正则取出书名,和文件名一对一记录。
最后再通过批量更名工具,通过刚才的一对一记录,一次性改名就行了。(很多更名软件比如支持)
我都试过, 包括质谱,千问.
- 这里面只有 deepseek 最聪明, 其他都不行. 会胡编, 等我发现时, 被豆包害惨了.
- 其他AI最多都支持10个文件上传,只有deepseek支持50个.
升级电脑吧,顺便买个
电脑的配置如何?
本地部署大模型的话,可能吗?不需要那些几十几百B的模型,Qwen3-VL-8B的模型处理这个问题足够了,甚至4B的模型一样可以
大致思路就是
- 通过工具,把PDF的一页或者几页拆分出去,并且转换成图片
- 把图片抛给大模型,让大模型给这个PDF起名字
- 之后把大模型给出的结果进行解析,我的意思是,这种需求,大概率的说大模型的输出会是JSON数据,得解析一下,如果通过自然语言返回的话,可能不稳定,至少我之前类似的项目,就不太稳定
- 最后,把文件重命名
- 然后开始处理下一个文件
不清楚你的电脑速度如何,我之前用类似的流程处理,我的4070显卡大约是15秒钟处理一张照片。虽然处理五万张一样要很久,但是至少比手动处理要好得多。
这里本地大模型的使用肯定有问题,处理完一个pdf后重置session甚至处理完一章就重置。这类任务不需要跨pdf的上下文,一口气都塞进去反而影响效果。
然后绝对不要用openclaw,这玩意的内置提示词做的稀烂,非常浪费token,而且最有价值的记忆之类的功能在你这个任务里又用不上,所以用opencode之类的编程agent搭一个pipeline就行。
如果只是需要根据内容改名便于之后检索,可以先用minerU之类的工具解析pdf,然后把解析结果取一小部分出来交给本地LLM判断是否是乱码,如果不是乱码就直接保存文本,如果是乱码就渲染成图片再用ocr处理(可能只需要处理个目录之类的就能总结内容?)。别指望ocr的结果完全正确,但大概正确的内容LLM总是可以大概理解的。最后甭管是解析的还是识别的都交给LLM总结,你可以先手动总结几个pdf并且将输入和结果写进system prompt作为示例学习的内容。
上述流程可以串行也可以并行,而且估计快不起来。我之前参与过一个类似的项目的大概就是上面说的这个流程,百万级别的pdf(单个十多页,大部分可以直接解析)用四五台八卡GPU的集群跑了两个多星期才完事。
如果都是出版物或者国家标准之类的文件,我的建议是全删了或者不去管他,然后有需要再一本本在其他站点下载,每次下载就命名好……
如果是不是烂大街资源,有很多内部资料,那当我没说……
5w本书,话说一个图书馆平均有多少本书,平均又有多少个图书管理员?
(也就说说平均一个图书管理员 管理多少本书?)
你这个工作量一个人完不成的 即使借助ai只能得到不太好的分类
的确都是出版物或国家标准这类文件
但是吧, 首先, 标准规范是有版权的, 不是能随便下载的.
其次, 很多出版物根本没有电子版, 能有扫描版已经属实不易了. 比如 50年代~90年代的出版物.
8B不行的,我试过总结小说库,一塌糊涂,主要是表现非常不稳定。
在可用、复读、驴唇不对马嘴见横跳,需要32B模型才能有稳定可用的调用结果
哦,是什么问题?
能说一下技术细节吗?比方说
- 模型的具体型号
- 系统题词是怎么写的
- 用户题词是怎么拼接的
- 不稳定的表现是什么
- 32B的模型又是哪个
回复部门: 标准创新管理司 时间:2024-06-14
《中华人民共和国标准化法》第十七条规定,“强制性标准文本应当免费向社会公开。国家推动免费向社会公开推荐性标准文本”。国家标准委发布的《推进国家标准公开工作实施方案》中要求, 强制性国家标准实现免费向社会公开 **;
非采标的推荐性国家标准实现免费向社会公开。采标的推荐性国家标准,在遵守国际(国外)标准组织版权政策前提下,实现免费向社会公开。
标准文本公开包括在线阅读和全文下载等方式,因涉及标准版权等因素,国家标准全文公开系统收录的现行有效强制性国家标准中非采标的可在线阅读和下载,采标的只可在线阅读;
系统收录的现行有效推荐性国家标准当前仅提供在线阅读服务,不提供全文下载服务。
目前未涉及版权保护的国家标准文本由国家标准委统一在官方平台——国家标准全文公开系统上免费公开(网址:https://openstd.samr.gov.cn/bzgk/gb/index)。
涉及版权保护的标准文本,可在中国标准服务网购买由中国标准出版社等相关单位出版的标准文本。**
- 模型的具体型号:Qwen3 8B
- 系统题词是怎么写的:
摘要提示词
### 指令:
你是一个专业的文学摘要生成器。请根据以下小说文本片段,生成一段简洁连贯的摘要。
### 要求:
1. 摘要需忠实反映原文内容、关键情节和人物关系
2. 语言精炼,逻辑清晰,只返回摘要内容
3. 严格禁止添加任何引导语、总结语或说明文字
4. 不要提及"要求"或"指令"
5. 不要包含"本段描述"、"这部分讲的是"等引导词
6. 不要包含"很高兴帮到您"、"以下是摘要"等无关内容
7. 字数控制在600字以内
### 文本片段:
{text}
### 摘要:
标签提示词
请仔细阅读以下小说文本片段,提取最能代表其内容特点的关键词作为标签。
要求:
1. 只返回具体的内容标签,不要返回描写手法标签(如不要'动作描写''心理描写')
2. 标签应体现题材、核心情节、人物特征、世界观、风格等核心元素
3. 标签要具体明确:用'科幻'代替'科幻小说',用'爱情'代替'爱情故事'
4. 标签数量不超过10个,用中文顿号"、"分隔
5. 不要任何解释和说明文字
文本片段:
{text}
标签:
-
不稳定的表现是什么:
不听话呀,一半的时候模型正常给出总结,一半的时候模型给出的结果会在大约200字后开始复读,比如:
正常总结200字……云端的一、一片、一片、一片、一片、一片、一片、一片、一片、一片、一片、一片……(循环)
正常总结150字,复读上边的150字,复读上边的150字……
前边总结是符合文章的,后面开始根据前边他自己写的总结继续瞎编剧情,而是不是总结小说 -
32B的模型又是哪个: deepseek r1 32b
尝试很多方案,比如提示词中设置明确的禁止项,不要凑字数,给出正面示例。比如参数上降低温度值,使用核采样,提升重复惩罚,都没明显的效果,提示词总有不听的,重复惩罚太低复读,太高就瞎说。
先不说别的差距,就说这个上下文窗口:
Qwen3 8B 上下文窗口只有32K,正常哪有顶格用的,基本每次只能输入1~1.5万字。正常网文小说一章就4K~6K字了,为了完整性,每次也就能输入2~3个章节。一本连载小说写个几十章,40万字不多吧。
模型没法直接吃一本进去,就只能切分,分别总结,再依次递归总结,直到最后的总结。
- 切40个分段,每个分段得到500字总结 = 2万字阶段总结
- 切2个分段,每个分段5000字总结=1万字总结 (最后两次需要尽可能保留信息)
- 得到最终的500~600字总结。
32B模型,上下文就到了128K了(可用4万字),一次可以吃更多的文字进去,递归总结次数就少了,质量和稳定性就高了
- 切10个分段,每个分段得到4000字总结 = 一共4万字阶段总结
- 得到最终的500~600字总结。
哦哦,确实,我用Qwen3 8B的时候也会遇到复读机的问题,甚至现在的Qwen3.5 9B也会有这个问题。当时折腾了好久,最后是限定以给出的JSON的格式来回答,才算是稳定。
不过Qwen3-VL-8B-Instruct没有遇到过类似的问题。而且,楼主的主力诉求是做OCR或者说图片的内容识别,VL 8B处理这个问题应该是可以的。
最后,我印象中Qwen3 8B没有Instruct的微调模型,也就是说,确实不太适合处理你所描述的这类内容。
我感觉小模型,主要不是能力,而是输入长度问题……
小说动则几十万上百万字,小模型只能拆分再组合,递归多次质量下降的饿太快,而且我也不知道为什么貌似AI生成的文字再让AI总结质量就会下滑,所以之前还有过long模型这种小参量,大上下文的流派。
上下文和注意力在大模型里就是一个问题
- 并不是模型能够接受的上下文多就更强,那个更多是设置,能够支持1B的上下文并不意味着它就更强
- 不是模型的上下文支持得多,就能够真的处理更多的内容。目前阶段,所有的模型都会在上下文增加之后,出现能力下降的情况。或许未来的模型能够从技术上解决这个问题,但是,目前的模型不行
- 绝大多数的模型在训练的时候,使用的是16K-64K的上下文素材,所以在这个范围内的回答会是高质量的。我们把一个128K的内容丢给他,大模型回答出来一堆垃圾是完全可以理解的
最后,让AI总结对话我这里没看到什么问题啊,我跟Gemini聊上半个小时,最后让他总结一下,形成一篇文档,他都做得挺好的啊…
是的是的
Qwen3 8B这种模型,理论上他的32K窗口等价于3~4W个汉字(1 个中文字符 ≈ 0.6 个 token)了,
但实际上一次最多也只能输入1~1.5万字,单轮输入输出加一起不超过窗口的一半,才能稳定得到可控的回答。
Gemini那玩意理论上窗口有1M(1024K)上下文窗口,和最长64K的输出。
汇报一下.
web服务效果很好, 但是费用比较高,而且需要手工拖动文件. 5w多个的文件, 每次200个, 也需要手工操作很多次.
所以这个当做备用方案了.
目前采用了 python , 配合 千问 qwen3-vl-flash 的 batch 操作.
采用的节省方案:
- 图片尺寸缩小,节省输入token
- batch模式,非实时处理, 价格便宜一半
- 只返回json, 后期再处理, 输出token用量节省
而且由于使用了 python跑代码, 只要保证全程开机即可, 不用繁琐的关注进度和手动操作.
代码我分享如下:
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""
并发批处理版(半价)+ 图片压缩 + 最多 5 个任务同时运行
- 每次最多同时提交 5 个批处理任务,避免串行等待
- 动态调度:一个任务完成后立即下载结果,并提交新任务
- 自动压缩图片(长边 672px),降低 token 消耗
- 每个批次完成后一次性更新总 JSON,防止结果丢失
- 输出 JSON 记录原始->新文件名映射,失败文件移入“已处理”目录
- 临时文件(batch_requests_*.jsonl, batch_results_*.jsonl)不会被删除,便于手工合并或调试
"""
import os
import json
import base64
import shutil
import time
import re
import requests
from PIL import Image
from threading import Lock
# ==================== 配置区 ====================
INPUT_PATH = r"D:\Down\提取" # 图片目录或单张图片
API_KEY = "" # 你的APIKEY
MODEL = "qwen3-vl-flash"
MAX_IMAGE_SIZE = 672 # 压缩长边
BATCH_SIZE = 50 # 每批处理的图片数量
MAX_CONCURRENT_BATCHES = 5 # 最大同时进行的批处理任务数
BATCH_COMPLETION_WINDOW = "24h" # 批处理完成等待时间
POLL_INTERVAL = 120 # 轮询间隔(秒)
# ===============================================
BASE_URL = "https://dashscope.aliyuncs.com/compatible-mode/v1"
HEADERS = {"Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json"}
SYSTEM_PROMPT = (
"判断类别:standard(图集/规范)或document(法规/其他)。去除发文机关。字段:\n"
"standard:code,name\n"
"document:title,docnum(无则\"\")\n"
"只返回JSON:{\"type\":\"standard或document\",...}无额外文字"
)
# 全局锁,用于保护 JSON 写入操作
json_lock = Lock()
# ------------------- 图像压缩 -------------------
def compress_image(image_path, max_size=MAX_IMAGE_SIZE):
"""压缩图片长边,返回压缩后的路径(若原图已满足则返回原路径)"""
with Image.open(image_path) as img:
if max(img.size) <= max_size:
return image_path
ratio = max_size / max(img.size)
new_size = (int(img.width * ratio), int(img.height * ratio))
resized = img.resize(new_size, Image.Resampling.LANCZOS)
temp_path = image_path + ".compressed.png"
resized.save(temp_path, optimize=True)
return temp_path
return image_path
def encode_image(image_path):
"""读取图片并返回 base64 编码(使用压缩后的路径)"""
with open(image_path, "rb") as f:
return base64.b64encode(f.read()).decode()
# ------------------- 批量 API 核心 -------------------
def prepare_jsonl_requests(png_files, output_jsonl_path, batch_index):
"""
生成单个批次的请求文件(JSONL)
返回:文件路径,以及 custom_id -> 原文件路径 的映射表
"""
mapping = {}
with open(output_jsonl_path, "w", encoding="utf-8") as f:
for idx, orig_path in enumerate(png_files):
compressed_path = compress_image(orig_path)
img_b64 = encode_image(compressed_path)
if compressed_path != orig_path:
os.remove(compressed_path)
custom_id = f"batch_{batch_index:04d}_req_{idx:06d}"
mapping[custom_id] = orig_path
body = {
"model": MODEL,
"messages": [
{"role": "system", "content": SYSTEM_PROMPT},
{
"role": "user",
"content": [
{"type": "text", "text": "请分析图片"},
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{img_b64}"}}
]
}
],
"temperature": 0.0,
"max_tokens": 200,
"response_format": {"type": "json_object"}
}
line = json.dumps(
{
"custom_id": custom_id,
"method": "POST",
"url": "/v1/chat/completions",
"body": body
},
ensure_ascii=False
)
f.write(line + "\n")
return output_jsonl_path, mapping
def upload_file(file_path, purpose="batch"):
"""上传 JSONL 文件,返回 file_id"""
url = f"{BASE_URL}/files"
headers_upload = {"Authorization": f"Bearer {API_KEY}"}
with open(file_path, "rb") as f:
files = {"file": (os.path.basename(file_path), f, "application/json")}
data = {"purpose": purpose}
resp = requests.post(url, headers=headers_upload, files=files, data=data)
resp.raise_for_status()
return resp.json()["id"]
def create_batch(input_file_id, endpoint="/v1/chat/completions", completion_window=BATCH_COMPLETION_WINDOW):
"""创建批处理任务,返回 task_id"""
url = f"{BASE_URL}/batches"
payload = {
"input_file_id": input_file_id,
"endpoint": endpoint,
"completion_window": completion_window
}
resp = requests.post(url, headers=HEADERS, json=payload)
resp.raise_for_status()
return resp.json()["id"]
def get_batch_status(task_id):
"""获取批处理任务状态,返回 (status, output_file_id, error_file_id)"""
url = f"{BASE_URL}/batches/{task_id}"
resp = requests.get(url, headers=HEADERS)
resp.raise_for_status()
data = resp.json()
return data.get("status"), data.get("output_file_id"), data.get("error_file_id")
def download_result(file_id, dest_path):
"""下载结果文件(JSONL)"""
url = f"{BASE_URL}/files/{file_id}/content"
resp = requests.get(url, headers=HEADERS)
resp.raise_for_status()
with open(dest_path, "w", encoding="utf-8") as f:
f.write(resp.text)
return dest_path
def sanitize(name):
"""清理非法文件名字符"""
for ch in r'<>:"/\|?*':
name = name.replace(ch, "")
name = name.strip().rstrip(".")
return name if name else "未命名"
def generate_filename(info):
"""根据 API 返回的 JSON 生成新文件名"""
t = info.get("type", "").strip().lower()
if t == "standard":
code = info.get("code", "")
name = info.get("name", "")
base = f"{code} {name}" if code and name else (code or name)
elif t == "document":
title = info.get("title", "")
docnum = info.get("docnum", "")
base = f"{title} ({docnum})" if docnum else title
else:
raise ValueError(f"未知类别: {t}")
base = sanitize(base)
if not base or base == "未命名":
raise ValueError("生成文件名为空")
return base + ".png"
def move_single_file(file_path, processed_dir):
"""移动单个文件到已处理目录,处理重名"""
if not os.path.exists(file_path):
return
target = os.path.join(processed_dir, os.path.basename(file_path))
if os.path.exists(target):
base, ext = os.path.splitext(os.path.basename(file_path))
counter = 1
while os.path.exists(os.path.join(processed_dir, f"{base}_{counter}{ext}")):
counter += 1
target = os.path.join(processed_dir, f"{base}_{counter}{ext}")
shutil.move(file_path, target)
def extract_json_from_text(text):
"""从文本中提取 JSON 对象"""
if not text:
return None
text = text.strip()
if text.startswith("```json"):
text = text[7:]
if text.startswith("```"):
text = text[3:]
if text.endswith("```"):
text = text[:-3]
text = text.strip()
try:
return json.loads(text)
except json.JSONDecodeError:
match = re.search(r'(\{.*\})', text, re.DOTALL)
if match:
try:
return json.loads(match.group(1))
except json.JSONDecodeError:
pass
return None
def parse_batch_result_and_update(result_file, mapping, output_json, processed_dir):
"""
解析单个批次的结果文件,将结果合并到总 JSON 中(加锁保护)
返回:更新后的结果字典(仅用于返回,实际已写入文件)
"""
# 加载已有结果
with json_lock:
if os.path.exists(output_json):
try:
with open(output_json, "r", encoding="utf-8") as f:
results = json.load(f)
except json.JSONDecodeError:
# 备份损坏文件
backup = output_json + ".broken.bak"
shutil.copy(output_json, backup)
print(f"警告: {output_json} 损坏,已备份至 {backup},将重新开始记录")
results = {}
else:
results = {}
newly_processed = 0
with open(result_file, "r", encoding="utf-8") as f:
for line in f:
if not line.strip():
continue
item = json.loads(line)
custom_id = item["custom_id"]
orig_path = mapping.get(custom_id)
if not orig_path:
continue
orig_name = os.path.basename(orig_path)
if orig_name in results:
continue # 已处理过
# 处理 API 返回的错误
error = item.get("error")
if error:
results[orig_name] = f"ERROR: {error.get('message', '未知错误')}"
move_single_file(orig_path, processed_dir)
newly_processed += 1
else:
body_resp = item.get("response", {}).get("body", {})
choices = body_resp.get("choices", [])
if not choices:
results[orig_name] = "ERROR: 无响应内容"
move_single_file(orig_path, processed_dir)
newly_processed += 1
continue
content = choices[0].get("message", {}).get("content", "")
info = extract_json_from_text(content)
if info is None:
results[orig_name] = f"ERROR: JSON解析失败,原始内容: {content[:200]}"
move_single_file(orig_path, processed_dir)
newly_processed += 1
continue
try:
new_name = generate_filename(info)
new_path = os.path.join(processed_dir, new_name)
counter = 1
base, ext = os.path.splitext(new_name)
while os.path.exists(new_path):
new_path = os.path.join(processed_dir, f"{base}_{counter}{ext}")
counter += 1
shutil.move(orig_path, new_path)
results[orig_name] = new_name
newly_processed += 1
except Exception as e:
results[orig_name] = f"ERROR: {e}"
move_single_file(orig_path, processed_dir)
newly_processed += 1
# 一次性写入更新后的结果
with json_lock:
with open(output_json, "w", encoding="utf-8") as jf:
json.dump(results, jf, ensure_ascii=False, indent=2)
print(f"批次解析完成,新增处理 {newly_processed} 张,当前总结果数: {len(results)}")
return results
def submit_batch(batch_pngs, batch_index, processed_dir):
"""
提交一个批处理任务(生成请求文件 -> 上传 -> 创建任务)
返回 task_info 字典(包含 task_id, mapping, jsonl_path 等)
若失败则返回 None,并将该批图片直接移入已处理目录
"""
print(f"\n===== 提交批次 {batch_index+1},共 {len(batch_pngs)} 张图片 =====")
jsonl_path = f"batch_requests_{batch_index:04d}.jsonl"
try:
print("生成批量请求文件(自动压缩图片)...")
_, mapping = prepare_jsonl_requests(batch_pngs, jsonl_path, batch_index)
print("上传请求文件...")
file_id = upload_file(jsonl_path)
print("创建批处理任务...")
task_id = create_batch(file_id)
print(f"任务创建成功,task_id: {task_id}")
return {
"task_id": task_id,
"batch_index": batch_index,
"jsonl_path": jsonl_path,
"mapping": mapping,
"file_id": file_id,
"status": "submitted"
}
except Exception as e:
print(f"提交批次 {batch_index+1} 失败: {e}")
# 失败时将该批次的所有图片直接移入已处理目录,记录错误
with json_lock:
if os.path.exists("rename_result.json"):
with open("rename_result.json", "r", encoding="utf-8") as f:
results = json.load(f)
else:
results = {}
for orig_path in batch_pngs:
orig_name = os.path.basename(orig_path)
if orig_name not in results:
results[orig_name] = f"ERROR: 提交失败 - {e}"
move_single_file(orig_path, processed_dir)
with open("rename_result.json", "w", encoding="utf-8") as f:
json.dump(results, f, ensure_ascii=False, indent=2)
# 注意:即使提交失败,我们也保留 jsonl_path 文件,便于手工排查
return None
def process_completed_task(task_info, output_json, processed_dir):
"""处理已完成的任务:下载结果、解析、更新总 JSON、保留临时文件"""
task_id = task_info["task_id"]
mapping = task_info["mapping"]
jsonl_path = task_info["jsonl_path"] # 保留此文件,不删除
print(f"\n任务 {task_id} 已完成,开始下载结果...")
try:
# 获取最终状态(确保拿到 output_file_id)
status, out_id, err_id = get_batch_status(task_id)
if status != "completed":
print(f"任务 {task_id} 状态为 {status},跳过结果处理")
if err_id:
err_file = f"batch_errors_{task_id}.jsonl"
download_result(err_id, err_file)
print(f"错误详情保存至 {err_file}")
# 将所有图片移入已处理目录并记录错误
with json_lock:
if os.path.exists(output_json):
with open(output_json, "r", encoding="utf-8") as f:
results = json.load(f)
else:
results = {}
for orig_path in mapping.values():
orig_name = os.path.basename(orig_path)
if orig_name not in results:
results[orig_name] = f"ERROR: 批处理未完成 - {status}"
move_single_file(orig_path, processed_dir)
with open(output_json, "w", encoding="utf-8") as f:
json.dump(results, f, ensure_ascii=False, indent=2)
return
result_file = f"batch_results_{task_id}.jsonl"
download_result(out_id, result_file)
print(f"结果已下载至 {result_file}")
print("解析结果并更新总 JSON...")
parse_batch_result_and_update(result_file, mapping, output_json, processed_dir)
# 注意:临时文件(jsonl_path 和 result_file)均被保留,不删除
# 以便在网络中断或需要手工合并时使用。您可以手动清理这些文件。
except Exception as e:
print(f"处理任务 {task_id} 结果时出错: {e}")
# 出错时将所有图片移入已处理目录并记录错误
with json_lock:
if os.path.exists(output_json):
with open(output_json, "r", encoding="utf-8") as f:
results = json.load(f)
else:
results = {}
for orig_path in mapping.values():
orig_name = os.path.basename(orig_path)
if orig_name not in results:
results[orig_name] = f"ERROR: 结果处理失败 - {e}"
move_single_file(orig_path, processed_dir)
with open(output_json, "w", encoding="utf-8") as f:
json.dump(results, f, ensure_ascii=False, indent=2)
# ------------------- 主流程(并发控制)-------------------
def main():
if not INPUT_PATH or INPUT_PATH == r"C:\Your\Image\Folder":
print("请先在代码中设置 INPUT_PATH 为实际图片目录路径!")
return
# 获取所有 PNG 文件列表
if os.path.isfile(INPUT_PATH):
if INPUT_PATH.lower().endswith(".png"):
png_files = [INPUT_PATH]
else:
print("指定文件不是 PNG 格式")
return
base_dir = os.path.dirname(INPUT_PATH)
elif os.path.isdir(INPUT_PATH):
png_files = []
for root, _, files in os.walk(INPUT_PATH):
for f in files:
if f.lower().endswith(".png"):
png_files.append(os.path.join(root, f))
base_dir = INPUT_PATH
else:
print("路径不存在")
return
if not png_files:
print("未找到任何 PNG 文件")
return
processed_dir = os.path.join(base_dir, "已处理")
os.makedirs(processed_dir, exist_ok=True)
output_json = "rename_result.json"
# 加载已有结果,跳过已处理的文件
if os.path.exists(output_json):
try:
with open(output_json, "r", encoding="utf-8") as f:
results = json.load(f)
print(f"已加载已有结果 {len(results)} 条")
except json.JSONDecodeError:
backup = output_json + ".broken.bak"
shutil.copy(output_json, backup)
print(f"警告: {output_json} 损坏,已备份至 {backup},将重新开始")
results = {}
else:
results = {}
processed_names = set(results.keys())
remaining = [fp for fp in png_files if os.path.basename(fp) not in processed_names]
print(f"总计 {len(png_files)} 张图片,已处理 {len(processed_names)} 张,剩余 {len(remaining)} 张")
if not remaining:
print("所有文件均已处理,无需调用 API")
return
# 将剩余图片分批
batches = []
idx = 0
for i in range(0, len(remaining), BATCH_SIZE):
batches.append(remaining[i:i + BATCH_SIZE])
total_batches = len(batches)
print(f"共分为 {total_batches} 个批次,每批最多 {BATCH_SIZE} 张")
print(f"最大并发任务数: {MAX_CONCURRENT_BATCHES}")
# 任务队列管理
active_tasks = [] # 每个元素为 task_info 字典
next_batch_index = 0 # 下一个待提交的批次索引
# 主循环
while True:
# 1. 提交新任务(如果还有未提交的批次且活跃任务数未达上限)
while len(active_tasks) < MAX_CONCURRENT_BATCHES and next_batch_index < total_batches:
batch_pngs = batches[next_batch_index]
task_info = submit_batch(batch_pngs, next_batch_index, processed_dir)
if task_info is not None:
active_tasks.append(task_info)
next_batch_index += 1
# 2. 如果没有活跃任务且所有批次都已提交,退出循环
if not active_tasks and next_batch_index >= total_batches:
break
# 3. 轮询所有活跃任务的状态
completed_tasks = []
for task in active_tasks:
try:
status, out_id, err_id = get_batch_status(task["task_id"])
task["status"] = status
if status in ["completed", "failed", "expired", "cancelled"]:
completed_tasks.append(task)
except Exception as e:
print(f"获取任务 {task['task_id']} 状态失败: {e}")
# 如果连续失败,可以标记为失败并移除(这里简单跳过,下次轮询再试)
# 4. 处理已完成的任务
for task in completed_tasks:
active_tasks.remove(task)
process_completed_task(task, output_json, processed_dir)
# 5. 如果没有任务完成且还有未提交任务,则等待一段时间再轮询
if not completed_tasks and (active_tasks or next_batch_index < total_batches):
time.sleep(POLL_INTERVAL)
print("\n全部批次处理完成!")
print("临时文件(batch_requests_*.jsonl 和 batch_results_*.jsonl)已保留,可手动删除。")
if __name__ == "__main__":
main()



