尧图网站设计 尧图网站设计YAOTU DESIGN
ARTICLE DETAIL

资讯详情

深耕网站设计与一线实操的经验洞察。

Agent错误分析

Agent错误分析 Muse -Agent 最危险的失败不一定来自“不会推理”而是来自“在不确定性尚未被消除时就执行了动作”。1 让muse 生成一个时政新闻和科技新闻 分别播报现象生成了两条语音进行播报时政新闻播报没有问题但是让播报科技新闻却一直在播报时政新闻而且每一段语音会在某一个词语处卡壳可能达到一个重复上限后才继续往后进行Personal Agent failureartifact identity resolution / referent grounding failure。这通常不是“模型不知道哪个视频内容更好”而是系统没有稳定地把自然语言里的“那个 A 的”“另一个 B 的”“第二个音频”绑定到唯一 artifact ID。Muse 官方确实把这类输出抽象成Artifacts而且它强调 artifact 可以脱离聊天文本、独立存在同时 Muse 还会维护长期 Goals、Memory 和 activity state。这意味着一旦同一个 conversation/task 里生成多个相似 artifact系统必须额外解决一个类似数据库主键的问题。User:Generate video for Event AGenerate video for Event B↓Muse task execution┌─────────┴─────────┐↓ ↓Artifact #1 Artifact #2typevideo typevideoeventA eventB↓ ↓video videosimilar metadata similar metadata↓User:Play the one for Event A↓Reference Resolver↓??? A or B ???为什么会 confused第一种可能是最基础的artifact metadata 不够 discriminative。因为是先生成的时政新闻 后生成的科技新闻两者的区分度问题导致每次让播报xx 新闻都去播报先生成的时政新闻 因为检索的top k就是按照时间来的后续可以验证一下[{type: video,title: Event video,created_at: ...},{type: video,title: Event video,created_at: ...}]第二个问题更深conversation memory 和 execution state 混在了一起。Agent State 管理问题很多 Agent 系统会把之前发生过什么重新塞进 contextUser asked Event A Generated video User asked Event B Generated video然后用户说Play the first one.模型通过语言推理猜first one probably Event A video这非常危险。正确做法不应该主要靠 LLM inference。应该有Conversation Memory ≠ Artifact Registry也就是Natural language memory User created videos for Event A and Event B Structured State artifact_A → id123 artifact_B → id456播放必须从 structured state 获取play(artifact_id123)而不是play(probably the Event A video)第三个问题叫referential ambiguityUncertainty问题其实包含很多方面。人类语言里面特别常见play that oneplay the other oneplay the previous oneplay Event As oneuse the voice file from before这些叫anaphora / coreference。LLM本身确实很擅长 coreference resolution但在 Agent 系统中linguistic coreference ≠ resource identity.比如模型可能有 90% 把握that video → Event A聊天的时候90%很好。但执行的时候delete file send email play private recording purchase something90% 完全不够。因此应该存在Natural Language ↓ Entity Resolution ↓ Artifact Resolution ↓ Confidence / ambiguity check ↓ Execution如果P(A)0.55 P(B)0.45不应该执行。应该询问Do you mean the Event A video or the Event B video?相关话题uncertainty reduction类型问题是什么例子Intent ambiguity用户到底想做什么“帮我处理一下这个”——是总结、播放还是发送Semantic/query ambiguity用户描述不足真实目标/实体不确定症状信息不充分多个疾病假设都可能Referential ambiguity用户说的“这个/那个/科技新闻”到底对应哪个具体对象“播放科技新闻”应该绑定tech_doc_02还是其他 artifactState-grounding error理解对了但执行状态仍指向旧对象intent播放科技新闻但 active_artifact 还是 politics_docTool-binding error前面都对最终 tool 参数传错play(politics_doc)而不是play(tech_doc)technology news↓应该解析到哪个 artifact↓tech_document_id ?所以这是更典型的referential grounding ambiguity / artifact-resolution failure而不是user intent ambiguity这本质是structured uncertainty reduction over candidate latent states需要semantic structural alignment这其实暴露了纯 embedding retrieval 的一个问题假设Politics news: US election, AI regulation, economy... Technology news: OpenAI, AI agents, semiconductor...用户说play the technology newsDense retrieval 可能因为两个 document 都出现AI US companies government technology导致 embedding similarity 很接近。例如sim(query, politics) 0.82 sim(query, technology) 0.85差距非常小。但加入结构约束type uploaded_news category technology status unplayed goal current_user_request以后Score α semantic_similarity β structural_alignment γ state_consistency δ recency/task relevance可能得到Politics: semantic .82 structure .25 state .10 Technology: semantic .85 structure .93 state .95于是对象解析就稳定很多。
返回列表