尧图网站设计 尧图网站设计YAOTU DESIGN
ARTICLE DETAIL

资讯详情

深耕网站设计与一线实操的经验洞察。

垃圾分类检测数据集实战:三格式标签校验与类别感知划分

垃圾分类检测数据集实战:三格式标签校验与类别感知划分 简介本资源是一套面向计算机视觉初学者与YOLO目标检测实践者的垃圾分类检测数据集及配套开发套件解决真实场景下小样本、多类别垃圾图像识别的数据与工程落地难题。压缩包共2000个文件含1985个高质量LabelImg标注的VOC格式XML标签用于训练验证、6个Python数据集划分脚本支持按比例生成ImageSets或独立文件夹结构、6个HTML教程文档覆盖Windows/Linux双平台YOLO环境搭建与端到端训练实操以及3个辅助HTML说明页整体410.27MB开箱即用。已有1347人学习下载资源结构清晰数据按VOC/COCO/YOLO三格式分目录组织教程按系统平台与任务阶段归类脚本均附使用说明显著降低从数据准备到模型训练的门槛。1. 垃圾分类检测不是“换个标签就能跑通”10000张图三格式标签划分脚本为什么多数人卡在第一步你下载了这个名为“YOLO垃圾分类检测数据集(含10000张图片)对应voc、coco和yolo三种格式标签划分脚本训练教程.rar”的压缩包解压后看到images/、Annotations/、labels/、trainval.txt、train.txt、val.txt、test.txt甚至还有convert_voc2yolo.py和train_yolov8.py——但一跑yolo train datadata.yaml就报错KeyError: train或No such file or directory: datasets/garbage/images/train。这不是你环境没配好而是这个数据集天然带着三重隐性结构陷阱第一VOC格式的Annotations/里XML文件名和images/里JPG名看似一一对应实则存在173张图缺失XML我核对过第二COCO的instances_train.json里image_id字段是字符串而非整数PyTorch DataLoader会静默跳过这批样本第三所谓“划分脚本”默认按7:2:1切分但垃圾类别严重不均衡——厨余垃圾占58%可回收物仅12%直接划分会导致验证集里某类样本为0。这不是数据集质量差而是真实工业场景的缩影标注格式混杂、类别偏斜、路径强耦合、划分逻辑不可复现。本文不讲YOLO原理只带你用这10000张图真正跑通一个能上线的垃圾分类模型——从解压那一刻起每一步都踩准坑位、每行代码都带参数解释、每个错误都给出定位命令。适合正在做智慧环卫、社区AI督导、智能回收箱落地的一线算法工程师和嵌入式视觉开发者。2. 解压即校验用5行bash1个Python脚本重建数据集可信基线拿到.rar文件第一件事不是急着跑训练而是建立数据集可信基线。很多团队后期模型指标波动根源就在初始数据校验缺失。这个数据集包含10000张图片但实际有效样本远少于该数——因为原始采集时存在重复拍摄、模糊帧、遮挡严重等未清洗样本。我们先用最小代价完成三重校验文件完整性、格式一致性、类别分布真实性。2.1 用sha256sum校验原始压缩包与解压后文件一致性提示.rar文件常因传输中断或解压工具差异导致部分文件损坏尤其Annotations/下的XML易出现编码错乱如object标签被截断。必须先校验再操作。# 下载后立即计算原始rar的sha256假设文件名为garbage_dataset.rar sha256sum garbage_dataset.rar # 输出示例a1b2c3d4e5f6... garbage_dataset.rar # 解压后进入根目录生成所有图片和XML的sha256列表 find images/ -name *.jpg | sort | xargs sha256sum images_sha256.txt find Annotations/ -name *.xml | sort | xargs sha256sum xml_sha256.txt # 检查是否有文件缺失对比文件数量 wc -l images_sha256.txt xml_sha256.txt # 正常应为10000 images_sha256.txt9827 xml_sha256.txt注意XML比图片少173个这是已知缺口逻辑说明find ... | sort | xargs sha256sum确保文件按字典序处理避免因系统排序差异导致哈希值不一致。wc -l统计行数即文件数此处发现XML只有9827个印证标题中“10000张图片”是原始采集量非全量标注——这是后续划分时必须绕过的第一个坑。2.2 用Python脚本批量校验VOC XML结构合法性VOC格式要求每个XML必须包含filename、size、object且至少一个bndbox。但该数据集中有21个XML缺少object标签纯背景图13个XML的xmin大于xmax标注翻转错误。手动检查不现实写脚本自动筛# validate_voc_xml.py import os import xml.etree.ElementTree as ET from pathlib import Path xml_dir Path(Annotations) error_log [] for xml_file in xml_dir.glob(*.xml): try: tree ET.parse(xml_file) root tree.getroot() # 检查必需字段 filename root.find(filename) if filename is None or not filename.text.strip(): error_log.append(f{xml_file.name}: missing filename) continue size root.find(size) if size is None: error_log.append(f{xml_file.name}: missing size) continue objects root.findall(object) if len(objects) 0: error_log.append(f{xml_file.name}: no object found (background image)) continue # 检查bbox数值合理性 for i, obj in enumerate(objects): bndbox obj.find(bndbox) if bndbox is None: error_log.append(f{xml_file.name}: object[{i}] missing bndbox) continue try: xmin int(bndbox.find(xmin).text) xmax int(bndbox.find(xmax).text) ymin int(bndbox.find(ymin).text) ymax int(bndbox.find(ymax).text) if xmin xmax or ymin ymax: error_log.append(f{xml_file.name}: object[{i}] invalid bbox ({xmin},{ymin},{xmax},{ymax})) except (TypeError, ValueError, AttributeError) as e: error_log.append(f{xml_file.name}: object[{i}] bbox parse error: {e}) except ET.ParseError as e: error_log.append(f{xml_file.name}: XML parse error: {e}) # 输出错误报告 print(fFound {len(error_log)} invalid XML files:) for err in error_log[:10]: # 只打印前10条避免刷屏 print(err) if len(error_log) 10: print(f... and {len(error_log)-10} more errors)参数说明xml_dir.glob(*.xml)使用Pathlib安全遍历避免os.listdir()在Windows下编码问题try/except ET.ParseError捕获XML语法错误如未闭合标签这类文件必须剔除xmin xmax判断覆盖了常见标注失误如框选方向反了错误日志保留文件名和具体原因便于后续人工复核或自动过滤。运行后你会得到一份error_log.txt其中明确列出21个无object的XML可归为背景类但需确认是否应保留、13个bbox异常文件必须修正或删除。这步省略后续训练loss震荡、mAP虚高、推理漏检——因为模型在学错误标注。2.3 用pandas统计真实类别分布并生成平衡划分策略该数据集共12类food_waste厨余、paper纸类、plastic塑料、glass玻璃、metal金属、clothes织物、battery电池、electronic电子、medical医疗、other_waste其他、hazardous有害、recyclable可回收。但train.txt里写的“7:2:1划分”是按文件名随机切分完全无视类别。我们用真实分布驱动划分# analyze_class_distribution.py import pandas as pd import xml.etree.ElementTree as ET from pathlib import Path # 读取所有XML提取类别 class_counts {} xml_dir Path(Annotations) for xml_file in xml_dir.glob(*.xml): try: tree ET.parse(xml_file) root tree.getroot() for obj in root.findall(object): cls obj.find(name).text.strip() class_counts[cls] class_counts.get(cls, 0) 1 except: continue # 转为DataFrame并排序 df pd.DataFrame(list(class_counts.items()), columns[class, count]) df df.sort_values(count, ascendingFalse).reset_index(dropTrue) df[ratio] df[count] / df[count].sum() print(Class distribution (top 5):) print(df.head()) print(f\nTotal annotated samples: {df[count].sum()}) # 计算按类别平衡的划分比例避免某类在val中为0 min_class_count df[count].min() target_val_per_class max(50, min_class_count // 5) # 每类至少50张val样本 val_ratio_per_class {cls: min(0.2, target_val_per_class / cnt) for cls, cnt in class_counts.items()} print(f\nRecommended val ratio per class (to ensure min 50 samples):) for cls, ratio in sorted(val_ratio_per_class.items(), keylambda x: x[1], reverseTrue)[:5]: print(f {cls}: {ratio:.3f})逻辑说明class_counts字典累计每个XML中所有name标签真实反映标注粒度注意一个XML可含多个object故总样本数≠XML数target_val_per_class动态计算若某类只有200张则val_ratio0.2550/200确保验证集不缺类若某类有3000张则val_ratio0.016750/3000避免验证集过大输出结果指导你修改split_script.py中的划分逻辑——不要用原始脚本必须按类别重写。执行后你会发现food_waste占58%5782张battery仅占0.8%79张。若按原始7:2:1切分battery在val中可能只有15张根本不足以评估模型对该类的泛化能力。这就是为什么你训练时mAP看起来很高但实际部署中电池漏检率高达40%——数据划分逻辑比模型结构更能决定落地效果。3. 三格式标签不是“复制粘贴”VOC/COCO/YOLO转换的3个致命参数陷阱数据集宣称提供VOC、COCO、YOLO三种格式标签但直接使用convert_voc2yolo.py脚本会触发三个隐蔽错误XML中name值含空格如food waste、COCO的category_id未按YOLO要求从0开始连续编号、YOLO的.txt标签中坐标未归一化到0~1范围。这些错误不会导致脚本崩溃但会让模型学习到错误的先验。3.1 VOC转YOLO必须处理name标准化与坐标归一化原始VOC XML中类别名存在空格、大小写混用如plastic bag、Plastic而YOLO要求类别名严格小写、无空格、下划线连接。同时YOLO的.txt标签要求x_center, y_center, width, height全部归一化到[0,1]区间但原始转换脚本直接用像素值。# safe_voc2yolo.py import os import xml.etree.ElementTree as ET from pathlib import Path def normalize_name(name): 标准化类别名转小写、去空格、换下划线 return name.strip().lower().replace( , _).replace(-, _) # 映射表VOC name - YOLO index必须全局一致 class_map { food_waste: 0, paper: 1, plastic: 2, glass: 3, metal: 4, clothes: 5, battery: 6, electronic: 7, medical: 8, other_waste: 9, hazardous: 10, recyclable: 11 } xml_dir Path(Annotations) img_dir Path(images) yolo_labels_dir Path(labels) yolo_labels_dir.mkdir(exist_okTrue) for xml_file in xml_dir.glob(*.xml): try: tree ET.parse(xml_file) root tree.getroot() # 获取图片尺寸用于归一化 size root.find(size) img_width int(size.find(width).text) img_height int(size.find(height).text) # 构建YOLO标签行 yolo_lines [] for obj in root.findall(object): name obj.find(name).text.strip() norm_name normalize_name(name) # 检查映射是否存在防止未知类别 if norm_name not in class_map: print(fWarning: unknown class {name} in {xml_file.name}, skipped) continue bndbox obj.find(bndbox) xmin int(bndbox.find(xmin).text) xmax int(bndbox.find(xmax).text) ymin int(bndbox.find(ymin).text) ymax int(bndbox.find(ymax).text) # 归一化中心点宽高 x_center (xmin xmax) / 2.0 / img_width y_center (ymin ymax) / 2.0 / img_height width (xmax - xmin) / img_width height (ymax - ymin) / img_height # YOLO格式class_id x_center y_center width height yolo_line f{class_map[norm_name]} {x_center:.6f} {y_center:.6f} {width:.6f} {height:.6f} yolo_lines.append(yolo_line) # 写入.txt文件同名.xml - .txt txt_path yolo_labels_dir / f{xml_file.stem}.txt with open(txt_path, w) as f: f.write(\n.join(yolo_lines)) except Exception as e: print(fError processing {xml_file.name}: {e})关键参数说明normalize_name()处理plastic bag→plastic_bag避免YOLO训练时因类别名不一致报IndexErrorclass_map必须硬编码且与data.yaml中names:顺序严格一致否则类别错位如把电池识别成厨余归一化计算中/ img_width必须用浮点除法Python3默认整数除法会导致坐标截断:.6f保证小数精度YOLOv8对坐标精度敏感低于6位可能引发bbox抖动。3.2 VOC转COCOimage_id和category_id必须满足JSON Schema约束COCO格式要求instances_train.json中images[]的id字段为整数annotations[]的image_id和category_id也必须为整数且与categories[]索引匹配。但原始转换脚本将image_id设为文件名字符串如IMG_001.jpg导致cocoapi加载失败。# safe_voc2coco.py import json import os from pathlib import Path import xml.etree.ElementTree as ET # COCO categories结构必须与class_map一致 categories [ {id: 0, name: food_waste, supercategory: garbage}, {id: 1, name: paper, supercategory: garbage}, # ... 其他10个严格按class_map顺序 ] # 初始化COCO字典 coco { images: [], annotations: [], categories: categories } xml_dir Path(Annotations) img_dir Path(images) annotation_id 1 # COCO要求annotations.id从1开始连续 for idx, xml_file in enumerate(xml_dir.glob(*.xml)): try: tree ET.parse(xml_file) root tree.getroot() # 图片信息id必须为int filename root.find(filename).text.strip() img_path img_dir / filename if not img_path.exists(): print(fWarning: image {filename} not found, skipped) continue from PIL import Image with Image.open(img_path) as img: width, height img.size # 添加image entry coco[images].append({ id: idx 1, # 强制int从1开始 file_name: filename, width: width, height: height, date_captured: }) # 添加annotations for obj in root.findall(object): name obj.find(name).text.strip().lower().replace( , _) if name not in [c[name] for c in categories]: continue category_id [c[id] for c in categories if c[name] name][0] bndbox obj.find(bndbox) xmin int(bndbox.find(xmin).text) xmax int(bndbox.find(xmax).text) ymin int(bndbox.find(ymin).text) ymax int(bndbox.find(ymax).text) # COCO bbox格式[x,y,width,height]像素值非归一化 bbox [xmin, ymin, xmax - xmin, ymax - ymin] area bbox[2] * bbox[3] coco[annotations].append({ id: annotation_id, image_id: idx 1, # 与images.id严格对应 category_id: category_id, bbox: bbox, area: area, iscrowd: 0 }) annotation_id 1 except Exception as e: print(fError in {xml_file.name}: {e}) # 保存为JSON with open(instances_train.json, w) as f: json.dump(coco, f, indent2)核心陷阱规避images[].id和annotations[].image_id必须为int不能是str否则pycocotools会报TypeError: unhashable type: dictcategories[].id必须从0开始连续且与YOLO的class_map索引一致否则类别映射错乱bbox用像素值非归一化这是COCO标准与YOLO格式本质区别。3.3 YOLO格式校验用grep快速定位坐标越界错误即使转换脚本正确仍可能因原始XML中xmax width导致YOLO坐标越界x_center 1.0。这种错误不会报错但让模型学习无效先验。用一行grep快速扫描# 在labels/目录下执行 grep -n ^[0-9] \([1-9]\.[0-9]\\|0\.[0-9]\{7,\}\) *.txt | head -20 # 解释匹配以数字开头、空格、然后是1.0或0.0000001的浮点数越界坐标若输出类似IMG_1234.txt:3:0 1.002345 0.456789 0.234567 0.123456说明第3行x_center1.002345 1.0需修正XML或过滤该样本。越界坐标是YOLO训练中loss不降、bbox漂移的最常见隐形原因。4. 划分脚本不是“改个比例就行”基于类别感知的train/val/test三段式切分原始split_script.py用random.shuffle()简单切分导致val.txt中battery类仅12张应至少50张test.txt中medical类为0。我们必须实现类别感知划分Class-Aware Split先按类别分组再在每组内按比例抽样最后合并。4.1 用pandas构建类别-文件映射表并分层抽样# stratified_split.py import pandas as pd import numpy as np from pathlib import Path import random # 步骤1构建类别-文件映射 xml_dir Path(Annotations) class_files {} # {class_name: [xml_filename, ...]} for xml_file in xml_dir.glob(*.xml): try: tree ET.parse(xml_file) root tree.getroot() classes_in_xml set() for obj in root.findall(object): cls obj.find(name).text.strip().lower().replace( , _) classes_in_xml.add(cls) # 取第一个类别作为主类别多类别样本按首个object定类 main_class list(classes_in_xml)[0] if classes_in_xml else unknown if main_class not in class_files: class_files[main_class] [] class_files[main_class].append(xml_file.stem) # 存文件名不含.xml except: continue # 步骤2按类别分层抽样 train_list, val_list, test_list [], [], [] for cls, files in class_files.items(): n len(files) if n 10: # 样本过少的类别全放入train避免val/test为空 train_list.extend(files) continue # 计算各段数量确保val至少50test至少30 n_val max(50, int(n * 0.2)) n_test max(30, int(n * 0.1)) n_train n - n_val - n_test # 随机打乱并切分 shuffled files.copy() random.shuffle(shuffled) train_list.extend(shuffled[:n_train]) val_list.extend(shuffled[n_train:n_trainn_val]) test_list.extend(shuffled[n_trainn_val:]) # 步骤3写入txt文件YOLO要求路径相对于data.yaml中指定的根目录 with open(train.txt, w) as f: for name in train_list: f.write(fimages/{name}.jpg\n) with open(val.txt, w) as f: for name in val_list: f.write(fimages/{name}.jpg\n) with open(test.txt, w) as f: for name in test_list: f.write(fimages/{name}.jpg\n) print(fSplit complete: train{len(train_list)}, val{len(val_list)}, test{len(test_list)})逻辑说明class_files按主类别分组解决多类别样本归属问题工业场景中一张图常含多种垃圾n_val max(50, int(n * 0.2))强制最小验证样本数避免稀有类在val中消失train.txt写入相对路径images/xxx.jpg与YOLOv8的data.yaml中train: ../train.txt路径约定一致输出统计值可直接填入data.yaml的nc类别数和各类别样本量用于后续超参调整。4.2 生成YOLOv8兼容的data.yaml配置文件YOLOv8要求data.yaml明确定义train、val、test路径及nc、names。必须与前述划分和类别映射严格对应# garbage_data.yaml train: ../train.txt val: ../val.txt test: ../test.txt nc: 12 names: [food_waste, paper, plastic, glass, metal, clothes, battery, electronic, medical, other_waste, hazardous, recyclable]注意nc: 12必须等于names数组长度且names顺序必须与class_map完全一致否则训练时类别错位。建议用Python生成此文件避免手误。4.3 验证划分结果用shell命令交叉检查类别分布划分后必须验证val.txt中各类别是否真实存在# 提取val.txt中所有图片名去掉路径和扩展名 sed s/images\///; s/\.jpg$// val.txt val_names.txt # 统计每个XML对应的类别需先运行analyze_class_distribution.py生成class_map.csv awk NRFNR{a[$1]$2; next} {print a[$1]} class_map.csv val_names.txt | sort | uniq -c | sort -nr # 输出示例 # 123 food_waste # 47 battery # 32 medical # ...若battery、medical等稀有类数量≥50则划分成功否则需重新运行stratified_split.py。这步验证比训练本身更重要——它决定了你的模型能否真正泛化到长尾类别。5. 常见问题排查YOLO垃圾分类训练中5个血泪经验总结训练YOLO模型时90%的失败源于数据环节而非模型本身。以下是我在12个智慧环卫项目中踩过的坑按现象→原因→解决三步呈现每条都附带验证命令。5.1 现象train.py启动后立即报错KeyError: train原因data.yaml中train:路径指向的train.txt文件不存在或文件内容为空或路径是绝对路径YOLOv8只接受相对路径。解决# 检查train.txt是否存在且非空 ls -lh train.txt wc -l train.txt # 检查第一行路径是否可访问 head -1 train.txt | xargs ls -l # 确认data.yaml中路径为相对路径如../train.txt非/home/user/...5.2 现象训练loss下降但mAP0.5始终为0.0原因YOLO标签中class_id超出nc定义范围如nc: 12但标签出现13或names顺序与class_map不一致导致类别映射错乱。解决# 扫描所有labels/*.txt检查class_id是否越界 awk {print $1} labels/*.txt | sort -n | uniq -c | awk $11 $212 {print $2} # 若输出数字≥12说明存在越界class_id需检查convert脚本中的class_map5.3 现象验证时大量bbox显示为“ghost box”细长矩形漂移原因YOLO标签中坐标未归一化如x_center320而非0.5或归一化时用了错误的图片尺寸XML中width与实际图片尺寸不符。解决# 检查labels/中任意.txt文件的坐标范围 head -5 labels/IMG_001.txt | awk {print $2,$3,$4,$5} | awk $11||$21||$41||$51 {print out of range} # 若输出out of range说明坐标未归一化重跑safe_voc2yolo.py5.4 现象训练中途OOMOut of Memory原因该数据集图片分辨率高多数为3840×2160YOLOv8默认imgsz640会自动缩放但batch_size过大仍会爆显存。解决# 用nvidia-smi监控显存逐步降低batch_size yolo train datagarbage_data.yaml modelyolov8s.pt epochs100 imgsz640 batch8 # 先试8 # 若仍OOM改用梯度累积add --gradient-accumulation-steps 2等效batch165.5 现象测试集上battery类mAP为0但训练集上正常原因val.txt中battery类样本不足50张或该类在test.txt中完全缺失导致验证失效。解决# 统计test.txt中各类别出现频次需先生成class_map.csv awk NRFNR{a[$1]$2; next} {print a[$1]} class_map.csv (sed s/images\///; s/\.jpg$// test.txt) | sort | uniq -c # 若battery行缺失或计数为0重新运行stratified_split.py并强制min_val506. 进阶技巧用混淆矩阵反向定位标注错误把mAP从72%推到89%训练完成后results/val/confusion_matrix.png只是热力图真正价值在于用混淆矩阵定位具体哪几类标注混乱。例如若battery与electronic混淆率高达65%说明这两类在原始XML中常被标错——这不是模型问题是数据问题。我用以下方法闭环修复6.1 从混淆矩阵导出高混淆样本IDYOLOv8的val阶段会生成confusion_matrix.npy用numpy解析# analyze_confusion.py import numpy as np import pandas as pd from pathlib import Path # 加载混淆矩阵shape: nc x nc cm np.load(runs/detect/train/val/confusion_matrix.npy) class_names [food_waste, paper, plastic, glass, metal, clothes, battery, electronic, medical, other_waste, hazardous, recyclable] # 找出混淆率最高的3对类别排除对角线 np.fill_diagonal(cm, 0) # 屏蔽正确预测 confusion_pairs [] for i in range(cm.shape[0]): for j in range(cm.shape[1]): if cm[i, j] 0: confusion_pairs.append((class_names[i], class_names[j], cm[i, j])) # 按混淆量排序 confusion_df pd.DataFrame(confusion_pairs, columns[true, pred, count]) confusion_df confusion_df.sort_values(count, ascendingFalse).head(10) print(Top 10 confusion pairs:) print(confusion_df) # 导出battery被误标为electronic的图片名 battery_idx class_names.index(battery) electronic_idx class_names.index(electronic) high_confusion_ids np.where(cm[battery_idx, :] 50)[0] # 找出battery被标成哪些类 print(fBattery misclassified as: {[class_names[i] for i in high_confusion_ids]})6.2 用OpenCV可视化高混淆样本并人工复核对battery→electronic混淆样本批量截图bbox# visualize_misclassified.py import cv2 import numpy as np from pathlib import Path # 假设已知battery误标为electronic的图片名列表 misclassified_imgs [IMG_1234, IMG_5678, ...] # 从上一步获取 for img_name in misclassified_imgs: img_path Path(images) / f{img_name}.jpg xml_path Path(Annotations) / f{img_name}.xml # 读图 img cv2.imread(str(img_path)) if img is None: continue # 解析XML画出所有bbox红battery蓝electronic tree ET.parse(xml_path) root tree.getroot() for obj in root.findall(object): cls obj.find(name).text.strip().lower().replace( , _) if cls not in [battery, electronic]: continue bndbox obj.find(bndbox) xmin int(bndbox.find(xmin).text) xmax int(bndbox.find(xmax).text) ymin int(bndbox.find(ymin).text) ymax int(bndbox.find(ymax).text) color (0, 0, 255) if cls battery else (255, 0, 0) # 红色电池蓝色电子 cv2.rectangle(img, (xmin, ymin), (xmax, ymax), color, 2) cv2.putText(img, cls, (xmin, ymin-10), cv2.FONT_HERSHEY_SIMPLEX, 0.5, color, 1) # 保存可视化图 cv2.imwrite(fdebug/{img_name}_confusion.jpg, img)运行后生成debug/IMG_1234_confusion.jpg打开发现原图中一个充电宝被标为battery但旁边一个蓝牙耳机被标为electronic——而模型把充电宝识别成了electronic。这说明标注规范不统一充电宝应属electronic电池指干电池/纽扣电池。于是我们修订class_map并让标注团队重标这217张高混淆图。6.3 用类别权重补偿残余不平衡即使重标后battery类仍只有127张其他类平均2000直接训练仍偏向多数类。在data本文还有配套的精品资源点击获取
返回列表