
1. 这不是“温度监控软件”而是内核级热管理中枢为什么你写的驱动一上电就烫手而高通平台却能稳压运行三小时你有没有遇到过这样的情况在嵌入式板子上跑一个简单的视频解码 demo不到两分钟 CPU 温度就飙到 95℃风扇狂转系统开始降频卡顿最后直接 thermal shutdown —— 板子黑屏重启。你查 dmesg只看到一行冷冰冰的thermal thermal_zone0: critical temperature reached, shutting down你翻设备树发现 thermal-zones 节点里一堆trip-point、cooling-maps、thermal-sensors但每个字段像天书你 grep 内核源码drivers/thermal/目录下二十多个 c 文件thermal_core.c三千行thermal_sysfs.c又两千行根本不知道从哪下手。这不是硬件散热设计的问题这是你没真正理解 Linux 内核 thermal framework 的通用架构。我做过六个不同 SoC 平台的 thermal 移植从全志 H3/H6 到瑞芯微 RK3399/RK3566从 NXP i.MX8MQ 到高通 QCS605再到 Intel Atom x5-E3930 工业主板。所有平台只要涉及温控策略定制、多传感器协同、动态功耗墙调节、甚至风扇曲线精细化控制最终都绕不开 thermal framework 这套机制。它不是可有可无的附加模块而是内核功耗子系统Power Management Subsystem中与 cpuidle、cpufreq、devfreq 并列的四大支柱之一。它的核心价值是把“温度”这个物理量抽象成内核可感知、可调度、可响应的第一类资源对象—— 就像内存有 struct pageCPU 有 struct cpumaskthermal zone 就是 struct thermal_zone_device它有自己的生命周期、状态机、事件通知链和策略引擎。关键词“Linux 内核功耗子系统”“thermal framework”“通用架构”不是空泛术语。它们指向一个事实thermal framework 是内核为所有芯片厂商提供的标准化热管理接口层。高通不用自己重写整套温控逻辑只需注册一个struct thermal_zone_device_ops实现读取 sensor 值、设置 trip 点瑞芯微不用重复实现 cooling device 绑定逻辑只需提供struct thermal_cooling_device_ops定义如何调频、如何关核、如何控风扇而你写的用户态 daemon比如 thermald只需要通过 sysfs 或 netlink 与这套框架交互无需关心底层 sensor 是 I2C 还是 ADC是单点还是阵列。这种分层解耦正是“通用架构”的本质——它不解决“怎么测温度”而是定义“温度数据如何被内核消费”它不规定“风扇该转多快”而是提供“冷却能力如何被调度”。如果你正在调试一个烫手的嵌入式项目或者准备面试 Linux 内核岗尤其涉及功耗/thermal/SoC 驱动方向又或者正要为国产芯片适配新内核版本那么这篇梳理不是理论科普而是你明天就要用的实操地图。它不讲“thermal 是什么”而是告诉你thermal_zone_device_register()注册时传的tzp参数里polling_delay和polling_delay_ps的区别到底在哪为什么trip_point[0]必须是 critical 类型而trip_point[1]却可以是 passivecooling_map中cdev_instance的weight字段是如何影响多 cooling device 协同降温的权重分配的这些细节文档不会写但代码会说话而我踩过的坑会替你把答案标出来。2. 架构全景图thermal framework 不是“一层皮”而是五层精密咬合的齿轮组很多人误以为 thermal framework 就是/sys/class/thermal/下一堆目录和文件或者thermal_zone_device_register()一个函数调用。这就像说汽车只是四个轮子加个方向盘。真正的 thermal framework 是一个纵向贯穿内核、横向连接各子系统的五层架构每一层都承担不可替代的职责且层间耦合极紧。我把它比作一台精密钟表最外层是表盘用户接口中间是擒纵机构核心调度再往里是游丝摆轮事件驱动最核心是发条盒传感器与执行器抽象而底座则是机芯基板内核基础设施。下面逐层拆解重点讲清每层“为什么这样设计”以及你在实际移植或调试中最容易卡死的位置。2.1 第一层硬件抽象层HAL—— sensor 与 cooling device 的统一建模这是整个框架的地基也是芯片厂商最先接触的部分。它不关心 sensor 是热敏电阻、红外测温芯片还是 SoC 内部的 PMIC ADC也不在意 cooling device 是 CPU 频率调节器、GPU 电压控制器还是 PWM 风扇驱动。它只做一件事把物理世界映射成内核可操作的对象。thermal_sensor通过struct thermal_sensor抽象但内核并未强制要求你实现这个结构体。更常见的是直接在struct thermal_zone_device中嵌入 sensor 读取逻辑。关键在于ops-get_temp()回调函数——它必须返回毫摄氏度m°C单位的整数值。我见过太多驱动在这里翻车有的返回 raw ADC 值有的除以 1000 后丢精度有的甚至用浮点运算内核禁止。正确做法是ret read_adc() * 1000 / scale_factor;确保单位严格对齐。cooling_device通过thermal_cooling_device_register()注册核心是struct thermal_cooling_device_ops。这里有两个致命陷阱get_max_state()和get_cur_state()返回的是state index0~max不是实际频率或 PWM 占空比set_cur_state()的输入 state 必须是离散的、预定义的档位如 0off, 125%, 250%, 375%, 4100%而非连续值。很多新手试图传入 37 这样的百分比值结果set_cur_state()直接静默失败——因为内核根本不校验只当它是合法索引去查表。提示HAL 层的健壮性直接决定上层是否可用。建议在get_temp()中加入超时检测和多次采样取中值逻辑避免单次异常读数触发误 shutdown在set_cur_state()中加入硬件确认机制如读回寄存器值比对防止“下发了但没生效”的幽灵问题。2.2 第二层热区抽象层Thermal Zone—— 温度域的逻辑容器与状态中心这是 thermal framework 的心脏。一个struct thermal_zone_device不是一个传感器而是一个温度管理域。它可以聚合多个 sensor如 CPU core temp GPU die temp PCB ambient temp也可以绑定多个 cooling device如 CPU freq GPU volt FAN PWM形成一个闭环控制单元。它的设计哲学是“温度”不是孤立的点而是需要综合评估的区域状态。Trip Point 设计struct thermal_trip数组是 zone 的决策中枢。每个 trip point 包含typecritical/passive/active、temp触发阈值单位 m°C、hysteresis迟滞值防止抖动。关键规则type THERMAL_TRIP_CRITICAL必须存在且temp最高触发后内核立即执行emergency_shutdown()不可中断、不可忽略type THERMAL_TRIP_PASSIVE触发后启动被动降温如调低 CPU 频率可通过thermal_zone_device_enable()/disable()动态开关type THERMAL_TRIP_ACTIVE触发后启动主动降温如开风扇需硬件支持 active cooling device。我移植 RK3399 时原厂 dts 把 GPU trip 设置为 passive但实际风扇是 active device。结果 GPU 温度一到阈值内核拼命降频风扇却纹丝不动——因为 passive trip 只调用cpufreqcooling device不触碰fan。修正方法要么改 trip type 为 active要么在 cooling map 中显式绑定 fan 到 passive trip。Polling 机制polling_delay毫秒和polling_delay_ps微秒的区别常被忽视。前者用于普通轮询如 I2C sensor后者用于高精度、低延迟场景如 SoC 内部快速 ADC。若polling_delay_ps 0内核会启用高精度定时器hrtimer否则用普通 timer。错误配置会导致 polling 频率远低于预期——比如设polling_delay1000你以为是 1s 一次但实际因 timer jitter 可能变成 1.2s错过关键升温拐点。2.3 第三层冷却映射层Cooling Map—— 多设备协同的权重调度引擎这是 thermal framework 最精妙也最容易被误解的一层。它不决定“哪个设备该工作”而是定义“当某个 trip 触发时哪些 cooling device 参与、以什么比例参与”。struct thermal_cooling_device_instance就是这张映射表的单元格。Weight 字段的真相weight不是百分比而是相对权重比。假设有两个 cooling devicecpu-freqweight10和gpu-freqweight5当 passive trip 触发需降温 100 units 时cpu-freq承担10/(105)*100 ≈ 66.7unitsgpu-freq承担5/(105)*100 ≈ 33.3units。这个计算由thermal_zone_trip_update()在update_temperature()中完成完全在内核态执行无用户态干预。Instance 绑定逻辑一个 cooling device 可以绑定到多个 trip point如 fan 绑定到 active 和 critical一个 trip point 也可以绑定多个 cooling device如 passive trip 同时绑定 cpu-freq 和 gpu-freq。但绑定关系必须在thermal_zone_device_register()前通过thermal_zone_bind_cooling_device()显式建立。遗漏绑定 该设备永远不会被调度这是调试中最常见的“设备注册了但不起作用”的原因。注意cooling map 的动态更新能力有限。thermal_zone_unbind_cooling_device()存在但实际使用极少。生产环境应一次性配准所有可能组合避免运行时反复绑定/解绑引发竞态。2.4 第四层事件分发与策略层Governor Notifier—— 内核态的决策大脑这一层负责“何时响应”和“如何响应”。它包含两个并行机制Governor温控策略内核内置三种 governorstep_wise阶梯式最常用、power_allocator基于 PID 的功率分配、bang_bang开关式仅用于 critical。step_wise的核心逻辑在step_wise_throttle()中根据当前温度与 trip threshold 的差值线性计算目标 cooling state。例如温度距 trip 点还有 5℃则 target_state max_state * (5 / hysteresis)。关键参数hysteresis必须合理设置——太小导致频繁切换风扇嗡嗡响太大导致响应迟钝温度冲高后才动作。Notifier Chain通知链当 zone 状态变化如进入/退出 trip、temperature 更新、device 绑定/解绑时内核会广播THERMAL_NOTIFY_*事件。用户态 thermald 或自定义 daemon 通过netlinksocket 监听这些事件。这是用户态深度介入的唯一合规通道。不要试图轮询 sysfs那是低效且易错的。2.5 第五层用户接口层Sysfs Netlink—— 人机交互的标准化窗口这是你每天打交道的地方但它的设计极具深意Sysfs 接口/sys/class/thermal/thermal_zone*/下的文件如temp、trip_point_0_temp、mode、policy。所有读写操作最终都落入thermal_sysfs.c的 file_operations。temp文件读取触发thermal_zone_get_temp()进而调用ops-get_temp()mode写入disabled会调用thermal_zone_device_disable()。Sysfs 不是“只读状态展示”而是实时控制总线。Netlink 接口NETLINK_THERMALfamily专为用户态 daemon 设计。相比 sysfs 轮询netlink 是事件驱动、零延迟、高可靠。thermald就是靠它实时获知THERMAL_TZ_TRIP_UP事件再决定是否启动更复杂的策略如记录日志、触发告警、调整 workload。这五层不是堆叠而是环环相扣HAL 提供原始数据 → Zone 聚合并判定状态 → Cooling Map 分配任务 → Governor 计算执行指令 → Sysfs/Netlink 对外暴露。任何一层断裂整个热管理就失效。而你的调试必须从 HAL 层开始逐层向上验证不能跳过。3. 核心实操从设备树到内核注册手把手复现一个可工作的 thermal zone光看架构不够得动手。下面以全志 H6 平台使用 sun50i-h6.dtsi为例完整走一遍 thermal zone 的构建流程。这不是照抄文档而是我实际移植时的 checklist包含所有隐藏坑点和验证步骤。3.1 设备树DTS定义语法正确 ≠ 功能正确DTS 是硬件描述的起点但内核解析 DTS 的逻辑非常苛刻。以下是一个典型 thermal zone 定义soc { thermal_zones: thermal-zones { cpu-thermal { thermal-sensors cpu_temperature; polling-delay 1000; // 单位ms polling-delay-passive 250; // 单位ms注意命名差异 trips { cpu_crit: cpu-crit { temperature 105000; // 105℃单位 m°C hysteresis 2000; // 2℃ type critical; }; cpu_alert: cpu-alert { temperature 85000; // 85℃ hysteresis 1000; // 1℃ type passive; }; }; cooling-maps { map0 { trip cpu_alert; cooling-device cpu0 0 0; // 绑定到 cpu0 的 freq cooling device }; map1 { trip cpu_alert; cooling-device gpu0 0 0; // 同时绑定 gpu0 }; }; }; }; };关键验证点极易出错thermal-sensors引用的cpu_temperature必须已正确定义且其compatible字符串必须匹配内核中thermal_of_sensor_register()所支持的 driver。H6 的 sensor driver 是sun50i-h6-ths所以cpu_temperature的compatible allwinner,sun50i-h6-ths必须一字不差。polling-delay-passive是 legacy 名称新内核5.10已弃用应使用polling-delay-ps微秒。但 H6 SDK 常用旧版内核故保留。混淆两者会导致 polling 完全失效。cooling-device cpu0 0 0中的0 0是min_state和max_state表示使用该 cooling device 的全部档位。若设为cpu0 1 3则只允许使用 state 1~3state 0最低频和 state 4最高频被禁用。trips节点名如cpu_crit必须唯一且不能与其它 zone 冲突。我曾因复制粘贴导致两个 zone 都叫cpu-crit内核解析时静默覆盖第二个 zone 的 trip 完全丢失。3.2 内核驱动注册三步缺一不可的初始化序列DTS 只是蓝图真正干活的是驱动。drivers/thermal/thermal_core.c中的thermal_zone_of_sensor_register()是入口但它背后有三步硬性依赖Step 1Sensor Driver 加载与注册cpu_temperature对应的 driver如sun50i_h6_ths.c必须在probe()中调用thermal_sensor_register()传入struct thermal_sensor实现ops-get_temp()确保返回值为 m°C 整数必须调用thermal_sensor_set_trips()如果 sensor 支持硬件 trip否则内核不会为其创建 zone。Step 2Thermal Zone 初始化thermal_of_add_tz()函数解析 DTS关键动作of_thermal_zone_init()为每个 zone 分配struct thermal_zone_deviceof_thermal_zone_bind()解析cooling-maps调用thermal_zone_bind_cooling_device()建立绑定of_thermal_zone_set_mode()根据mode属性默认enabled决定是否启动 polling。Step 3Cooling Device 注册与绑定cpu0和gpu0的 cooling device 必须在 zone 初始化前注册完毕。cpufreq_cooling_register()是标准接口但要注意cpufreq_cooling_register()返回的struct thermal_cooling_device*必须保存后续thermal_zone_bind_cooling_device()需要用到绑定顺序无关但注册顺序必须早于 zone 初始化。若 cooling device 注册晚于 zoneof_thermal_zone_bind()会找不到 device绑定失败且无报错。实操心得在dmesg中搜索thermal正常流程应看到[ 1.234567] thermal: thermal zone cpu-thermal registered [ 1.234589] thermal: bound cooling device cpu0 to thermal zone cpu-thermal [ 1.234612] thermal: bound cooling device gpu0 to thermal zone cpu-thermal若缺失bound行一定是 cooling device 注册时机或名称不匹配。3.3 Sysfs 验证用最朴素的方法确认每一层都在工作不要依赖cat /sys/class/thermal/thermal_zone0/temp一次成功就认为 ok。必须分层验证HAL 层验证# 查看 sensor 是否被识别 cat /sys/class/thermal/thermal_zone0/type # 应输出 cpu-thermal cat /sys/class/thermal/thermal_zone0/temp # 读取温度多次执行看是否变化 # 如果返回 0 或 -ENODEV说明 sensor driver 未加载或 get_temp() 返回错误Zone 层验证# 查看 trip points 是否加载 ls /sys/class/thermal/thermal_zone0/trip_point* cat /sys/class/thermal/thermal_zone0/trip_point_0_temp # 应为 105000 cat /sys/class/thermal/thermal_zone0/trip_point_1_temp # 应为 85000 # 修改 mode 测试 echo disabled /sys/class/thermal/thermal_zone0/mode cat /sys/class/thermal/thermal_zone0/mode # 应输出 disabled echo enabled /sys/class/thermal/thermal_zone0/modeCooling Map 验证# 查看绑定的 cooling device ls /sys/class/thermal/thermal_zone0/cdev* # 输出类似cdev0 cdev1对应 cpu0 和 gpu0 cat /sys/class/thermal/thermal_zone0/cdev0/cur_state # 当前 state cat /sys/class/thermal/thermal_zone0/cdev0/max_state # 最大 state # 手动触发 cooling测试 cooling device 是否真能工作 echo 3 /sys/class/thermal/thermal_zone0/cdev0/cur_state # 然后用硬件工具如逻辑分析仪确认 CPU 频率是否真的降到第 3 档Governor 验证# 查看当前策略 cat /sys/class/thermal/thermal_zone0/policy # 默认 step_wise # 强制切换策略仅测试用 echo power_allocator /sys/class/thermal/thermal_zone0/policy # 观察 temperature 变化时 cur_state 是否平滑变化而非阶梯跳变这四步验证每一步失败都指向不同层级的问题。它比任何 debug print 都高效因为它是内核框架自身提供的健康检查协议。4. 真实排障实录六个让我熬夜到凌晨的 thermal 问题与根因分析理论再完美不如一个真实 bug 有说服力。以下是我在六个项目中遇到的典型 thermal 问题附带 root cause、排查路径和永久修复方案。这些问题90% 的开发者都会先怀疑硬件或用户态但最终都指向 thermal framework 的某处隐性约束。4.1 问题一温度读数恒为 0dmesg 无报错sensor driver probe 成功现象cat /sys/class/thermal/thermal_zone0/temp始终返回 0dmesg | grep thermal显示 “registered”但无错误。排查路径cat /sys/class/thermal/thermal_zone0/type确认 zone 名正确ls /sys/class/thermal/thermal_zone0/确认temp文件存在strace cat /sys/class/thermal/thermal_zone0/temp发现read()系统调用返回 0进入 kernel debugprintk在thermal_zone_get_temp()中发现tz-ops-get_temp()返回 -EAGAINRoot Causesensor driver 的get_temp()实现中ADC 采样后未等待转换完成即读取寄存器返回 raw 值 0。内核将 0 解释为 0m°C而非错误。修复方案在get_temp()中加入readl_poll_timeout()等待 ADC ready bit或增加重试逻辑最多 3 次每次间隔 10us。4.2 问题二trip point 触发后cooling device 无响应cur_state不变现象温度升至 85℃dmesg打印thermal thermal_zone0: Trip point cpu-alert reached但cat /sys/class/thermal/thermal_zone0/cdev0/cur_state仍为 0。排查路径ls /sys/class/thermal/thermal_zone0/cdev*确认 cdev0 存在cat /sys/class/thermal/thermal_zone0/cdev0/cur_state和max_state确认 cooling device 自身工作正常cat /sys/class/thermal/thermal_zone0/mode确认是enabledcat /sys/class/thermal/thermal_zone0/policy确认是step_wiseRoot Causecooling-maps中cooling-device cpu0 0 0的0 0被解析为min_state0, max_state0即只允许 state 0。而 state 0 通常定义为 “最高频”所以cur_state永远是 0无法降低。修复方案DTS 中改为cpu0 0 4假设 max_state 是 4或在 driver 中确保get_max_state()返回正确的最大档位数。4.3 问题三系统在 70℃ 就 shutdown远低于 critical trip 的 105℃现象dmesg显示critical temperature reached但cat /sys/class/thermal/thermal_zone0/temp读数只有 7000070℃。排查路径cat /sys/class/thermal/thermal_zone0/trip_point_0_temp确认是 105000dmesg | grep -A 5 -B 5 critical查看前后文发现thermal thermal_zone1: critical temperature reached—— 原来是另一个 zone如 gpu-thermal触发了。Root Cause系统有多个 thermal zonecpu/gpu/pcb但dmesg日志只显示 zone name不显示温度值。thermal_zone1的 critical trip 被设为 70℃且其 sensor 读数异常偏高。修复方案ls /sys/class/thermal/列出所有 zone逐一检查temp和trip_point_0_temp用thermal_zone_device_disable()临时禁用可疑 zone 进行隔离。4.4 问题四风扇启动后无法停止即使温度已回落至 trip 下限现象温度升至 active trip60℃风扇启动cur_state3温度降至 55℃hysteresis5℃风扇仍保持cur_state3。排查路径cat /sys/class/thermal/thermal_zone0/trip_point_2_hysteresis确认 hysteresis 是 5000cat /sys/class/thermal/thermal_zone0/trip_point_2_temp确认是 60000cat /sys/class/thermal/thermal_zone0/temp确认当前是 55000Root Causestep_wisegovernor 的hysteresis仅用于 trip 进入/退出判断不用于 cooling state 的回退。它只会根据当前温度与 trip 的差值计算 target state而不会记忆“上次 state 是多少”。当温度从 61℃ 降到 55℃差值从 1℃ 变为 5℃target state 仍可能是 3。修复方案改用bang_banggovernor仅适用于 on/off 设备或在用户态 thermald 中实现带 hysteresis 的 state 回退逻辑。4.5 问题五thermal zone 注册失败dmesg 报Failed to register thermal zone现象内核启动时dmesg显示Failed to register thermal zone cpu-thermal无更多线索。排查路径grep -r Failed to register drivers/thermal/定位到thermal_zone_device_register()中的if (!tz)tz分配失败原因是kzalloc(sizeof(*tz), GFP_KERNEL)返回 NULL检查内存cat /proc/meminfo | grep MemFree发现启动时 MemFree 1MBRoot Causethermal zone 结构体较大约 2KB在内存紧张的嵌入式系统尤其是 initramfs 阶段可能分配失败。GFP_KERNEL在 low memory 时会 sleep但 early boot 时不可 sleep。修复方案在thermal_zone_device_register()前将GFP_KERNEL替换为GFP_ATOMIC或确保系统启动时预留足够内存修改memkernel param。4.6 问题六用户态 thermald 无法收到 netlink 事件轮询 sysfs 导致 CPU 占用 100%现象thermald进程 CPU 占用率持续 100%strace显示疯狂poll()sysfs 文件。排查路径ss -uln | grep 3000确认 thermald 是否监听 netlinkcat /proc/$(pidof thermald)/fdinfo/* | grep netlink确认 fd 是 netlinktcpdump -i lo -n -A port 3000无任何包Root Causethermald编译时未启用NETLINK支持configure flag--enable-netlink或内核 config 中CONFIG_NETLINK_MMAP未开启导致 netlink socket 创建失败thermald自动 fallback 到 sysfs 轮询。修复方案重新编译thermald确保./configure --enable-netlink检查内核.config确认CONFIG_NETLINK_MMAPy。这些问题每一个都曾让我在凌晨三点盯着示波器波形和 dmesg 日志发呆。它们共同揭示了一个事实thermal framework 的强大恰恰源于它的严谨。它不宽容任何模糊地带——单位必须是 m°Cstate 必须是离散索引binding 必须在注册前完成。理解这些“不宽容”就是掌握它的开始。5. 进阶实践超越默认配置用 thermal framework 实现工业级温控策略当你已经能让 thermal framework 正常工作下一步就是让它为你所用。默认的step_wise和 sysfs 接口只够应付消费电子。在工业控制、车载系统、边缘 AI 服务器等场景你需要更精细、更鲁棒、更可审计的温控策略。以下是三个经过量产验证的进阶方案全部基于 thermal framework 原生能力无需 patch 内核。5.1 方案一多 zone 协同的“温度场均衡”策略需求背景某边缘 AI 盒子搭载双 NPUNPU0/NPU1各自有独立 thermal zone。但散热风道设计导致 NPU0 温度常比 NPU1 高 10℃长期运行后 NPU0 性能衰减更快。实现原理利用 thermal framework 的THERMAL_NOTIFY_TEMP事件和thermal_zone_get_temp()API在用户态 daemon 中实现跨 zone 温度补偿。核心代码逻辑// 监听所有 zone 的温度事件 netlink_socket socket(AF_NETLINK, SOCK_RAW, NETLINK_THERMAL); bind(netlink_socket, addr, sizeof(addr)); while (1) { recv(netlink_socket, buf, sizeof(buf), 0); if (event THERMAL_NOTIFY_TEMP zone_name npu0) { int temp_npu0, temp_npu1; read_temp(npu0, temp_npu0); // 读取 npu0 温度 read_temp(npu1, temp_npu1); // 读取 npu1 温度 int delta temp_npu0 - temp_npu1; if (delta 5000) { // 超过 5℃ // 主动降低 npu0 的 cooling state提升 npu1 的 cooling state set_cooling_state(npu0-cdev, get_cur_state(npu0-cdev) - 1); set_cooling_state(npu1-cdev, get_cur_state(npu1-cdev) 1); } } }优势不修改内核完全在用户态实现策略可动态调整如根据负载类型切换均衡强度日志可审计记录每次补偿的 delta 和 state 变化。5.2 方案二基于历史数据的“预测性降频”策略需求背景某车载信息娱乐系统视频解码负载具有强周期性每 30s 一个 peak。传统 thermal 响应总是滞后导致每次 peak 都触发降频卡顿。实现原理thermal framework 的 polling 机制提供稳定的时间基准。用户态 daemon 记录过去 N 个 polling 周期的温度变化率ΔT/Δt当预测未来 1s 内温度将突破 trip point 时提前触发 cooling。关键参数polling_delay 5000.5s 一次保证采样密度使用滑动窗口window size 10计算平均升温速率预测模型predicted_temp current_temp rate * prediction_horizonprediction_horizon 10001s实施效果在 peak 到来前 300ms 开始降频避免了卡顿且平均温度比传统策略低 3℃。5.3 方案三安全关键场景的“双路冗余温控”需求背景某工业 PLC 控制器要求 thermal shutdown 必须 100% 可靠。不能依赖单一 kernel thread 或单一 sensor。实现原理利用 thermal framework 的criticaltrip