{"schema":"https://fuuraa.com/schemas/ai-evidence-atlas/v1","name":{"en":"FUURAA AI Evidence Atlas","zh":"FUURAA™ 赋睐™ AI 证据图谱"},"description":{"en":"A public, bilingual synthesis of consequential AI questions, primary evidence, competing interpretations, limitations and open problems.","zh":"面向重要 AI 问题、一手证据、不同解释、局限与未解问题的双语公共综合判断体系。"},"publisher":{"name":"FUURAA™","legalName":"FUURAA HOLDING GROUP PTE. LTD.","url":"https://fuuraa.com"},"licenceNote":"FUURAA original summaries and synthesis. Rights in linked external sources remain with their respective owners.","version":"1.0","generatedAt":"2026-07-28","recordCount":6,"records":[{"id":"persistent-agent-identity-and-memory","number":"01","version":"1.0","created":"2026-07-28","reviewed":"2026-07-28","field":{"en":"Agent foundations","zh":"Agent 基础"},"fieldPath":"/research/fields/ai-identity-memory","question":{"en":"Under what conditions can an AI Agent preserve identity and useful memory without silently expanding authority or privacy risk?","zh":"AI Agent 在什么条件下可以保持身份与有用记忆，同时不让权限或隐私风险被悄然扩大？"},"shortTitle":{"en":"Persistent agent identity and memory","zh":"持续性 Agent 身份与记忆"},"evidenceStatus":{"id":"developing","label":{"en":"Developing","zh":"证据正在积累"},"explanation":{"en":"The direction is visible, but the evidence base, operating history or independent validation remains incomplete.","zh":"方向已经可见，但证据基础、运行历史或独立验证仍不完整。"}},"currentEvidencePosition":{"en":"Standards for verifiable identity, delegated authority and selective disclosure are emerging, while research is showing that long-term memory requires temporal, relational and contradiction-aware reasoning. A complete, portable and accountable continuity layer has not yet been demonstrated across organisations and tools.","zh":"可验证身份、委托权限与选择性披露的标准正在形成；研究也显示，长期记忆需要时间感知、关系理解与矛盾处理能力。但跨组织、跨工具且可问责的完整连续性层，尚未得到充分证明。"},"scope":{"en":"This dossier examines technical and governance evidence for agent credentials, delegated authority, long-term memory, provenance, consent, revocation and cross-system continuity. It does not claim that FUURAA has completed such an infrastructure.","zh":"本档案研究 Agent 凭证、委托权限、长期记忆、来源证明、同意、撤销与跨系统连续性的技术及治理证据；不代表 FUURAA 已完成相关基础设施。"},"supportedClaims":[{"en":"Agent identity and delegated authority are distinct security problems from ordinary human login.","zh":"Agent 身份与委托权限，是不同于普通用户登录的独立安全问题。"},{"en":"Machine-verifiable credentials and selective disclosure can reduce reliance on centralised, over-collecting identity flows.","zh":"机器可验证凭证与选择性披露，可以降低对集中式、过度收集信息的身份流程的依赖。"},{"en":"Long-horizon memory quality depends on time, relationships, conflict handling and downstream decision usefulness—not only similarity search.","zh":"长期记忆质量取决于时间、关系、矛盾处理和对后续决策的价值，而不只是相似度检索。"},{"en":"Authority boundaries, connector controls, audit trails and revocation need to operate throughout the agent lifecycle.","zh":"权限边界、连接器控制、审计轨迹与撤销机制需要贯穿 Agent 全生命周期。"}],"evidenceTensions":[{"en":"More persistent context can improve usefulness while increasing privacy exposure, behavioural inference and the cost of an incorrect memory.","zh":"更持久的上下文可以提升实用性，也可能扩大隐私暴露、行为推断与错误记忆造成的代价。"},{"en":"Portable identity may reduce platform lock-in, yet portability can also move compromised authority or sensitive memory across boundaries.","zh":"可携带身份可能减少平台锁定，但也可能把被滥用的权限或敏感记忆带过组织边界。"},{"en":"A continuous agent may benefit from stable identity, while some contexts require deliberate separation, expiry or anonymity.","zh":"持续运行的 Agent 可能受益于稳定身份，但某些场景需要主动隔离、到期失效或匿名。"}],"openQuestions":[{"en":"Which identity and memory primitives will interoperate across vendors without creating a universal surveillance layer?","zh":"哪些身份与记忆原语能够跨供应商互操作，同时避免形成普遍监控层？"},{"en":"How should an agent explain which memories changed a consequential action?","zh":"Agent 应如何说明哪些记忆影响了一项重要行动？"},{"en":"What should survive when a user changes provider, role, jurisdiction or relationship to an organisation?","zh":"当用户更换服务商、角色、司法辖区或组织关系时，哪些身份与记忆应当继续保留？"}],"testableQuestions":[{"en":"Can an independent verifier reconstruct an agent’s authority chain and detect expired or excessive permissions?","zh":"独立核验方能否重建 Agent 的权限链，并识别已过期或过度授权？"},{"en":"Does contradiction-aware memory improve decisions on long-running tasks without materially increasing sensitive-data retention?","zh":"矛盾感知记忆能否改善长期任务决策，同时不显著增加敏感数据留存？"},{"en":"Can a person revoke a delegated capability and observe that revocation across every connected tool?","zh":"个人能否撤销一项委托能力，并确认撤销已在所有关联工具中生效？"}],"decisionRelevance":[{"en":"Research: evaluate memory through downstream decisions, provenance and correction—not recall alone.","zh":"科研：通过后续决策、来源证明与修正能力评估记忆，而不只测试召回率。"},{"en":"Engineering: separate identity, authority, memory and connector policy so each can be inspected and revoked.","zh":"工程：分离身份、权限、记忆与连接器策略，使每一层都可检查、可撤销。"},{"en":"Governance: require visible consent, purpose limitation, retention boundaries and incident-ready audit records.","zh":"治理：要求可见同意、目的限制、保存边界与可用于事故处置的审计记录。"}],"primaryEvidence":[{"id":"agents-need-first-class-identities","title":{"en":"Agents need first-class identities","zh":"Agent需要第一等数字身份"},"source":"NIST NCCoE","sourceTitle":"New Concept Paper on Identity and Authority of Software Agents","sourceUrl":"https://www.nist.gov/news-events/news/2026/02/new-concept-paper-identity-and-authority-software-agents","published":"2026-02-05","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"NIST's concept work applies identity standards and authorization practices directly to software and AI agents that access data, tools and applications.","zh":"NIST概念工作把身份标准与授权实践直接应用于访问数据、工具和应用的软件及AI Agent。"},"implication":{"en":"An agent should be identifiable independently from its user, model provider and runtime so responsibility can be traced precisely.","zh":"Agent应与用户、模型供应商和运行环境分别识别，才能准确追溯责任。"}},{"id":"delegated-authority-must-be-bounded","title":{"en":"Delegated authority must be explicit and bounded","zh":"委托权限必须明确且有边界"},"source":"NIST NCCoE","sourceTitle":"New Concept Paper on Identity and Authority of Software Agents","sourceUrl":"https://www.nist.gov/news-events/news/2026/02/new-concept-paper-identity-and-authority-software-agents","published":"2026-02-05","reviewed":"2026-07-26","status":"emerging","horizon":"1-3y","relation":"supports-or-limits","summary":{"en":"Agent autonomy creates a difference between authenticating an actor and deciding what that actor may do on another party's behalf.","zh":"Agent自治使“确认行动者身份”与“决定其可代表他人做什么”成为两个不同问题。"},"implication":{"en":"Future systems should encode scope, duration, resource limits, revocation and escalation into every delegation.","zh":"未来系统应把范围、期限、资源限制、撤销和升级机制写入每一次授权。"}},{"id":"selective-disclosure-can-reduce-agent-data-exposure","title":{"en":"Selective disclosure can reduce agent data exposure","zh":"选择性披露可减少Agent数据暴露"},"source":"W3C","sourceTitle":"Verifiable Credentials Data Model v2.0","sourceUrl":"https://www.w3.org/TR/vc-data-model-2.0/","published":"2025-05-15","reviewed":"2026-07-26","status":"emerging","horizon":"1-3y","relation":"supports-or-limits","summary":{"en":"Credential systems can be designed so a holder proves a relevant attribute without sharing an entire underlying identity record.","zh":"凭证系统可以让持有者证明相关属性，而无需共享完整底层身份记录。"},"implication":{"en":"Agents may demonstrate authority or eligibility while minimizing personal and enterprise information revealed to counterparties.","zh":"Agent可证明权限或资格，同时尽量减少向交易对手披露个人和企业信息。"}},{"id":"long-term-memory-needs-relation-awareness","title":{"en":"Long-term memory needs relation awareness","zh":"长期记忆需要理解记录之间的关系"},"source":"SubtleMemory research team","sourceTitle":"SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents","sourceUrl":"https://arxiv.org/abs/2606.05761","published":"2026-06-04","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"SubtleMemory tests whether agents distinguish complementary, nuanced and contradictory memories instead of retrieving isolated similar passages.","zh":"SubtleMemory测试Agent能否区分互补、细微差异和矛盾记忆，而不是只检索相似片段。"},"implication":{"en":"Memory stores should represent relationships among records, not only vector proximity.","zh":"记忆库应表达记录之间的关系，而不只是向量距离。"}},{"id":"contradictions-should-not-be-silently-merged","title":{"en":"Contradictions should not be silently merged","zh":"矛盾记忆不应被静默合并"},"source":"SubtleMemory research team","sourceTitle":"SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents","sourceUrl":"https://arxiv.org/abs/2606.05761","published":"2026-06-04","reviewed":"2026-07-26","status":"emerging","horizon":"1-3y","relation":"supports-or-limits","summary":{"en":"Persistent assistants accumulate records that may diverge as people, plans and circumstances change.","zh":"持续运行的助手会积累因人员、计划和环境变化而相互偏离的记录。"},"implication":{"en":"Memory systems should preserve versions, provenance and unresolved conflicts until a trusted process resolves them.","zh":"记忆系统应保留版本、来源和未解决冲突，直到可信流程完成裁定。"}},{"id":"connector-governance-becomes-an-enterprise-control-plane","title":{"en":"Connector governance becomes an enterprise control plane","zh":"连接器治理将成为企业控制平面"},"source":"OpenAI","sourceTitle":"Introducing AgentKit","sourceUrl":"https://openai.com/index/introducing-agentkit/","published":"2025-10-06","reviewed":"2026-07-26","status":"forecast","horizon":"1-3y","relation":"supports-or-limits","summary":{"en":"Central connector registries can determine which agents reach which systems and under what administrative policy.","zh":"集中式连接器目录可以决定哪些Agent能访问哪些系统，以及适用什么管理政策。"},"implication":{"en":"Organizations may manage agent connectivity like identity infrastructure: inventoried, approved, monitored and revocable.","zh":"组织可能像管理身份基础设施一样管理Agent连接：登记、批准、监测并可撤销。"}}],"relatedAnalysis":[{"id":"verifiable-credentials-agent-identity","path":"/news/insights/verifiable-credentials-agent-identity","title":{"en":"The Agent Era Will Need Machine-Verifiable Identity and Authority","zh":"Agent 时代需要机器可验证的身份与授权"},"sourceUrl":"https://www.w3.org/TR/vc-data-model-2.0/","reviewed":"27 July 2026"},{"id":"long-context-agent-memory","path":"/news/insights/long-context-agent-memory","title":{"en":"Long-Lived AI Requires More Than a Larger Context Window","zh":"长期存在的 AI，需要的不只是更大的上下文窗口"},"sourceUrl":"https://deepmind.google/research/publications/121073/","reviewed":"27 July 2026"},{"id":"ai-agent-standards","path":"/news/insights/ai-agent-standards","title":{"en":"Why AI Agents Need Infrastructure, Not Only Stronger Models","zh":"为什么 AI Agent 需要基础设施，而不只是更强的模型"},"sourceUrl":"https://www.nist.gov/news-events/news/2026/02/announcing-ai-agent-standards-initiative-interoperable-and-secure","reviewed":"27 July 2026"},{"id":"responsible-agentic-ai","path":"/news/insights/responsible-agentic-ai","title":{"en":"From Singapore: Foundations for Responsible Agentic AI","zh":"从新加坡出发：负责任部署 Agentic AI 的基础"},"sourceUrl":"https://www.imda.gov.sg/resources/press-releases-factsheets-and-speeches/factsheets/2026/updated-model-ai-governance-framework-for-agentic-ai","reviewed":"27 July 2026"}]},{"id":"multi-agent-coordination-and-interoperability","number":"02","version":"1.0","created":"2026-07-28","reviewed":"2026-07-28","field":{"en":"Agent systems","zh":"Agent 系统"},"fieldPath":"/research/fields/agent-systems","question":{"en":"When do multiple AI Agents outperform a well-designed single-agent system, and what coordination infrastructure makes that advantage dependable?","zh":"多个 AI Agent 在什么条件下能够优于设计良好的单 Agent 系统？哪些协同基础设施能让这种优势变得可靠？"},"shortTitle":{"en":"Multi-agent coordination and interoperability","zh":"多 Agent 协同与互操作"},"evidenceStatus":{"id":"mixed","label":{"en":"Mixed","zh":"证据与解释存在分歧"},"explanation":{"en":"Credible evidence supports more than one interpretation, or outcomes vary materially by context.","zh":"可信证据支持不止一种解释，或结果会随场景、群体与实施方式显著变化。"}},"currentEvidencePosition":{"en":"Open protocols, capability discovery and multi-agent research are advancing, but credible evidence also shows that adding agents can increase communication overhead, duplicated work and coordination failure. Benefits appear task-dependent and require explicit roles, shared state, verification and recovery.","zh":"开放协议、能力发现与多 Agent 研究正在进展；但可信证据也显示，增加 Agent 可能带来沟通开销、重复劳动与协同失败。收益高度依赖任务，并需要明确角色、共享状态、结果核验与恢复机制。"},"scope":{"en":"This dossier covers agent-to-agent protocols, tool interoperability, task allocation, shared state, commitment tracking, verification and failure recovery. It does not treat protocol support as proof that an agent or tool is trustworthy.","zh":"本档案覆盖 Agent 间协议、工具互操作、任务分配、共享状态、承诺跟踪、结果核验与失败恢复；不会把“支持协议”视为 Agent 或工具可信的证明。"},"supportedClaims":[{"en":"Capability discovery, declared interfaces and common tool protocols can reduce bespoke integration work.","zh":"能力发现、声明式接口与通用工具协议可以减少重复的定制集成工作。"},{"en":"Explicit task roles, commitment tracking and verifiable outputs are recurring requirements in effective multi-agent designs.","zh":"明确任务角色、跟踪承诺与核验输出，是有效多 Agent 设计中反复出现的要求。"},{"en":"Neutral stewardship can improve the prospects for cross-vendor protocols, but governance remains part of the technical problem.","zh":"中立治理有助于跨供应商协议发展，但治理本身仍是技术问题的一部分。"},{"en":"Coordination should be evaluated against a simpler baseline, including cost, latency, reliability and recovery.","zh":"协同系统应与更简单的基线进行比较，包括成本、延迟、可靠性与恢复能力。"}],"evidenceTensions":[{"en":"Specialisation can improve task coverage while increasing hand-off loss and shared-state complexity.","zh":"专业化可以扩大任务覆盖，也会增加交接损失与共享状态复杂度。"},{"en":"Open protocols improve portability, yet shared interfaces can widen the attack surface and propagate untrusted context.","zh":"开放协议提升可携带性，也可能扩大攻击面并传播不可信上下文。"},{"en":"Parallel work can reduce elapsed time, but coordination and verification may cost more than the work itself.","zh":"并行工作可能缩短时间，但协同与核验成本也可能超过任务本身。"}],"openQuestions":[{"en":"Which task characteristics reliably predict a multi-agent advantage?","zh":"哪些任务特征能够可靠预测多 Agent 的优势？"},{"en":"How should reputation, liability and incident recovery work across independently operated agents?","zh":"由不同主体运营的 Agent 之间，应如何处理信誉、责任与事故恢复？"},{"en":"What minimum shared semantics are needed before agents can negotiate commitments rather than merely exchange messages?","zh":"Agent 要从“交换消息”走向“协商承诺”，至少需要哪些共享语义？"}],"testableQuestions":[{"en":"Does the multi-agent design beat a single-agent baseline after communication, verification and recovery costs are included?","zh":"计入沟通、核验与恢复成本后，多 Agent 设计是否仍优于单 Agent 基线？"},{"en":"Can every output be traced to task ownership, evidence and the agent that accepted responsibility?","zh":"每项输出能否追溯到任务归属、证据以及接受责任的 Agent？"},{"en":"How does the system behave when one agent is unavailable, malicious or confidently wrong?","zh":"当一个 Agent 不可用、具有恶意或自信地出错时，系统会如何运行？"}],"decisionRelevance":[{"en":"Research: publish simple baselines and coordination overhead, not only best-case aggregate performance.","zh":"科研：公开简单基线与协同开销，而不只展示最佳聚合表现。"},{"en":"Engineering: make roles, state transitions, commitments, evidence and recovery paths observable.","zh":"工程：让角色、状态变化、承诺、证据与恢复路径保持可观察。"},{"en":"Procurement: protocol compatibility should not replace tool assurance, permission review or supplier accountability.","zh":"采购：协议兼容不能代替工具保证、权限审查与供应商问责。"}],"primaryEvidence":[{"id":"more-agents-do-not-guarantee-better-results","title":{"en":"More agents do not guarantee better results","zh":"更多Agent并不保证更好结果"},"source":"Stanford University and SAP Labs","sourceTitle":"CooperBench: Why Coding Agents Cannot be Your Teammates Yet","sourceUrl":"https://arxiv.org/abs/2601.13295","published":"2026-01-19","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"CooperBench found that paired coding agents performed worse on average than a single agent completing both tasks, exposing a coordination penalty.","zh":"CooperBench发现，成对协作的编程Agent平均表现低于单个Agent完成全部任务，暴露出协作成本。"},"implication":{"en":"Multi-agent architecture should be justified by measured division-of-labour gains, not by agent count.","zh":"多Agent架构应以可测量的分工收益为依据，而不是以Agent数量为依据。"}},{"id":"commitment-tracking-is-a-missing-capability","title":{"en":"Commitment tracking is a missing agent capability","zh":"承诺追踪是Agent团队缺失的能力"},"source":"Stanford University and SAP Labs","sourceTitle":"CooperBench: Why Coding Agents Cannot be Your Teammates Yet","sourceUrl":"https://arxiv.org/abs/2601.13295","published":"2026-01-19","reviewed":"2026-07-26","status":"emerging","horizon":"1-3y","relation":"supports-or-limits","summary":{"en":"The benchmark identifies vague communication, broken commitments and incorrect expectations about teammates as recurring failure modes.","zh":"该基准识别出模糊沟通、违背承诺和错误预期等反复出现的协作失败。"},"implication":{"en":"Agent teams need explicit shared state for promises, dependencies, conflicts and verification.","zh":"Agent团队需要明确共享承诺、依赖、冲突和验证状态。"}},{"id":"agent-to-agent-protocols-enter-neutral-governance","title":{"en":"Agent-to-agent protocols are entering neutral governance","zh":"Agent间协议正在进入中立治理"},"source":"Google Cloud and Linux Foundation","sourceTitle":"Google Cloud Donates A2A to Linux Foundation","sourceUrl":"https://developers.googleblog.com/google-cloud-donates-a2a-to-linux-foundation/","published":"2025-06-23","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"Google transferred the A2A specification and tooling to a Linux Foundation project supported by multiple enterprise technology companies.","zh":"Google把A2A规范和工具移交给Linux Foundation项目，并获得多家企业技术公司的支持。"},"implication":{"en":"Neutral stewardship can reduce vendor lock-in and make cross-company agent collaboration more credible.","zh":"中立治理可以减少供应商锁定，使跨企业Agent协作更可信。"}},{"id":"capability-discovery-precedes-agent-collaboration","title":{"en":"Capability discovery precedes agent collaboration","zh":"能力发现是Agent协作的前提"},"source":"Google Cloud and Linux Foundation","sourceTitle":"Google Cloud Donates A2A to Linux Foundation","sourceUrl":"https://developers.googleblog.com/google-cloud-donates-a2a-to-linux-foundation/","published":"2025-06-23","reviewed":"2026-07-26","status":"emerging","horizon":"1-3y","relation":"supports-or-limits","summary":{"en":"A2A is designed to help agents discover capabilities and exchange tasks without sharing their internal memory or implementation.","zh":"A2A旨在让Agent无需共享内部记忆或实现方式，也能发现能力并交换任务。"},"implication":{"en":"Future agent marketplaces will need trustworthy capability descriptions, versioning and verification.","zh":"未来Agent市场需要可信的能力描述、版本管理和验证机制。"}},{"id":"tool-connections-are-standardizing","title":{"en":"Tool connections are standardizing","zh":"工具连接正在标准化"},"source":"Model Context Protocol","sourceTitle":"Model Context Protocol Specification 2025-06-18","sourceUrl":"https://modelcontextprotocol.io/specification/2025-06-18/index","published":"2025-06-18","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"MCP defines a common protocol for connecting language-model applications with external data sources, resources and tools.","zh":"MCP定义了连接语言模型应用与外部数据源、资源及工具的通用协议。"},"implication":{"en":"Standard connectors can lower integration costs, but each connection still needs explicit trust and permission design.","zh":"标准连接器可以降低集成成本，但每个连接仍需明确设计信任与权限。"}},{"id":"protocol-support-does-not-equal-tool-trust","title":{"en":"Protocol support does not equal tool trust","zh":"支持协议不等于信任工具"},"source":"Model Context Protocol","sourceTitle":"Model Context Protocol Specification 2025-06-18","sourceUrl":"https://modelcontextprotocol.io/specification/2025-06-18/index","published":"2025-06-18","reviewed":"2026-07-26","status":"emerging","horizon":"now","relation":"supports-or-limits","summary":{"en":"A common connection format improves compatibility but cannot determine whether a server, tool output or requested action is safe.","zh":"通用连接格式提升兼容性，却不能判断服务器、工具输出或请求行动是否安全。"},"implication":{"en":"Registries, provenance, sandboxing and least-privilege controls must surround protocol adoption.","zh":"协议采用必须配套注册目录、来源链、沙箱和最小权限控制。"}}],"relatedAnalysis":[{"id":"multi-agent-coordination","path":"/news/insights/multi-agent-coordination","title":{"en":"The Multi-Agent Future Depends on Coordination, Not Just Quantity","zh":"多 Agent 的未来，关键不只是数量，而是协调能力"},"sourceUrl":"https://hai.stanford.edu/news/ai-coding-agents-fail-at-teamwork","reviewed":"27 July 2026"},{"id":"ai-agent-standards","path":"/news/insights/ai-agent-standards","title":{"en":"Why AI Agents Need Infrastructure, Not Only Stronger Models","zh":"为什么 AI Agent 需要基础设施，而不只是更强的模型"},"sourceUrl":"https://www.nist.gov/news-events/news/2026/02/announcing-ai-agent-standards-initiative-interoperable-and-secure","reviewed":"27 July 2026"},{"id":"responsible-agentic-ai","path":"/news/insights/responsible-agentic-ai","title":{"en":"From Singapore: Foundations for Responsible Agentic AI","zh":"从新加坡出发：负责任部署 Agentic AI 的基础"},"sourceUrl":"https://www.imda.gov.sg/resources/press-releases-factsheets-and-speeches/factsheets/2026/updated-model-ai-governance-framework-for-agentic-ai","reviewed":"27 July 2026"},{"id":"global-ai-index-2026","path":"/news/insights/global-ai-index-2026","title":{"en":"AI Adoption Is Accelerating, While Agent Deployment Remains Early","zh":"AI 采用正在加速，但 Agent 部署仍处早期"},"sourceUrl":"https://hai.stanford.edu/ai-index/2026-ai-index-report","reviewed":"27 July 2026"}]},{"id":"embodied-ai-safety-and-human-oversight","number":"03","version":"1.0","created":"2026-07-28","reviewed":"2026-07-28","field":{"en":"Embodied intelligence","zh":"具身智能"},"fieldPath":"/research/fields/embodied-intelligence","question":{"en":"What evidence is required before embodied AI can be trusted to act around people in open, changing environments?","zh":"具身 AI 在开放且不断变化的环境中与人共同活动之前，需要具备哪些证据才能获得信任？"},"shortTitle":{"en":"Embodied AI safety and human oversight","zh":"具身 AI 安全与人类监督"},"evidenceStatus":{"id":"developing","label":{"en":"Developing","zh":"证据正在积累"},"explanation":{"en":"The direction is visible, but the evidence base, operating history or independent validation remains incomplete.","zh":"方向已经可见，但证据基础、运行历史或独立验证仍不完整。"}},"currentEvidencePosition":{"en":"World models, vision-language-action systems and cross-embodiment learning are expanding robot capability, while renewed safety standards and monitoring practices are developing in parallel. Open-world reliability, rare-event behaviour and meaningful human intervention remain incompletely evidenced.","zh":"世界模型、视觉—语言—动作系统与跨形态学习正在扩展机器人能力，安全标准与监测实践也在同步发展。但开放世界可靠性、罕见事件行为与有效的人类介入，仍缺乏充分证据。"},"scope":{"en":"This dossier examines evidence on world models, robot foundation models, transfer, human oversight, monitoring, safety standards and deployment learning. Industry demonstrations are treated as bounded evidence, not proof of general autonomy.","zh":"本档案研究世界模型、机器人基础模型、迁移、人类监督、监测、安全标准与部署学习；行业演示只作为有边界的证据，不作为通用自主能力的证明。"},"supportedClaims":[{"en":"Predictive world representations and heterogeneous training data can improve transfer across tasks and embodiments.","zh":"预测性世界表征与异构训练数据可以改善跨任务、跨机器人形态的迁移。"},{"en":"Robot policies are increasingly combining perception, language, planning and low-level control in shared architectures.","zh":"机器人策略正在把感知、语言、规划与底层控制整合进共享架构。"},{"en":"Industrial and service-robot safety requires system-level controls beyond model evaluation alone.","zh":"工业与服务机器人安全需要系统级控制，不能只依赖模型评估。"},{"en":"Monitoring, operating limits and incident analysis need to continue after deployment.","zh":"监测、运行边界与事故分析需要在部署后持续进行。"}],"evidenceTensions":[{"en":"Learning from deployment can improve performance while exposing people and environments to immature behaviour.","zh":"从部署中学习可以提升性能，也可能让人和环境承受不成熟行为的风险。"},{"en":"Generalist policies may reduce retraining while making failure boundaries harder to specify.","zh":"通用策略可能减少重复训练，却也可能让失败边界更难描述。"},{"en":"Remote supervision can extend human reach, but attention limits and latency can make nominal oversight ineffective.","zh":"远程监督可以扩大人类能力范围，但注意力限制与延迟可能让名义上的监督失效。"}],"openQuestions":[{"en":"Which evaluations predict safe behaviour in homes, streets and workplaces rather than curated demonstrations?","zh":"哪些评估能够预测机器人在家庭、街道与工作场所的安全表现，而不是只适用于精心设计的演示？"},{"en":"How should responsibility be allocated among model provider, robot maker, deployer, operator and site owner?","zh":"模型提供方、机器人制造商、部署方、操作员与场所所有者之间，应如何分配责任？"},{"en":"What evidence should be required before a robot can expand its operating envelope?","zh":"机器人扩大运行范围之前，必须具备哪些证据？"}],"testableQuestions":[{"en":"Can the system detect and safely exit conditions outside its validated operating envelope?","zh":"系统能否识别超出已验证运行边界的情形，并安全退出？"},{"en":"Can a human supervisor understand, interrupt and recover the system within the time available?","zh":"人类监督者能否在可用时间内理解、打断并恢复系统？"},{"en":"Do safety results persist across sites, hardware variants, people and environmental change?","zh":"安全结果能否在不同场所、硬件版本、人群与环境变化中保持？"}],"decisionRelevance":[{"en":"Research: evaluate rare events, distribution shift and recovery—not only average task success.","zh":"科研：评估罕见事件、分布变化与恢复能力，而不只看平均任务成功率。"},{"en":"Engineering: combine model controls with hardware interlocks, safe states, observability and site design.","zh":"工程：把模型控制与硬件联锁、安全状态、可观察性及场所设计结合起来。"},{"en":"Governance: require deployment-specific evidence, named human responsibilities and a route for incident learning.","zh":"治理：要求与部署场景相匹配的证据、明确的人类责任与事故学习路径。"}],"primaryEvidence":[{"id":"interactive-world-models-move-beyond-video-generation","title":{"en":"World models are moving beyond passive video generation","zh":"世界模型正在超越被动的视频生成"},"source":"Google DeepMind","sourceTitle":"Genie 3: A new frontier for world models","sourceUrl":"https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/","published":"2025-08-05","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"Genie 3 demonstrates text-conditioned environments that can be navigated and changed in real time while preserving short-term consistency. This turns generated media into an interactive system with state and consequences.","zh":"Genie 3 展示了能够根据文字生成、实时导航和改变，并在一段时间内保持一致性的环境。这使生成媒体成为具有状态与后果的交互系统。"},"implication":{"en":"The frontier is not only more realistic imagery, but environments in which people and agents can act, observe outcomes and learn.","zh":"前沿不只是更逼真的画面，而是人类和 Agent 能够行动、观察结果并学习的环境。"}},{"id":"video-pretraining-teaches-predictive-world-structure","title":{"en":"Video pretraining can teach predictive structure about the world","zh":"视频预训练可以学习世界的预测性结构"},"source":"Meta AI Research","sourceTitle":"V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning","sourceUrl":"https://ai.meta.com/research/publications/v-jepa-2-self-supervised-video-models-enable-understanding-prediction-and-planning/","published":"2025-06-11","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"V-JEPA 2 learns from large-scale video and image data to represent motion, anticipate actions and support planning. It emphasizes prediction in a learned representation rather than reconstructing every visible pixel.","zh":"V-JEPA 2 从大规模视频和图像中学习运动表征、动作预判与规划能力，重点是在学习到的表示空间中预测，而不是重建每一个可见像素。"},"implication":{"en":"Observation-rich training may become a major route to physical understanding, especially where labelled robot demonstrations are scarce.","zh":"在标注机器人示范稀缺的领域，以观察为主的训练可能成为获得物理理解的重要路径。"}},{"id":"steerable-vlas-combine-what-and-how","title":{"en":"Robot policies are learning both what to do and how","zh":"机器人策略开始同时理解“做什么”和“怎么做”"},"source":"Physical Intelligence","sourceTitle":"π0.7: A Steerable Model with Emergent Capabilities","sourceUrl":"https://www.pi.website/blog/pi07","published":"2026-04-16","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"π0.7 accepts varied conditioning such as language, visual subgoals and performance metadata to steer task strategy as well as task identity.","zh":"π0.7接收语言、视觉子目标和性能元数据等多种条件，用于控制任务本身及执行策略。"},"implication":{"en":"Future robot interfaces may let operators specify speed, quality, caution and method without rewriting controllers.","zh":"未来操作员可直接指定速度、质量、谨慎程度与方法，而无需重写控制器。"}},{"id":"dual-speed-architectures-enter-humanoid-foundation-models","title":{"en":"Dual-speed architectures are entering humanoid models","zh":"双速度架构正在进入人形机器人模型"},"source":"NVIDIA","sourceTitle":"NVIDIA Announces Isaac GR00T N1 and Simulation Frameworks to Speed Robot Development","sourceUrl":"https://investor.nvidia.com/news/press-release-details/2025/NVIDIA-Announces-Isaac-GR00T-N1--the-Worlds-First-Open-Humanoid-Robot-Foundation-Model--and-Simulation-Frameworks-to-Speed-Robot-Development/default.aspx","published":"2025-03-18","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"GR00T N1 combines a slower vision-language reasoning system with a faster action model for continuous movement.","zh":"GR00T N1把较慢的视觉语言推理系统与较快的连续动作模型结合。"},"implication":{"en":"Humanoid stacks may increasingly separate deliberate planning from reflex-like control while testing their interaction as one safety system.","zh":"人形机器人将更多地分离深思规划与反射式控制，同时把二者交互作为一个整体安全系统测试。"}},{"id":"robot-safety-standards-are-being-renewed","title":{"en":"Industrial robot safety standards are being renewed","zh":"工业机器人安全标准正在更新"},"source":"ISO","sourceTitle":"ISO 10218-1:2025 — Robotics — Safety Requirements — Part 1: Industrial Robots","sourceUrl":"https://www.iso.org/standard/73933.html","published":"2025-02-05","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"The third edition of ISO 10218-1 updates safety requirements for industrial robots as machines before system integration.","zh":"ISO 10218-1第三版更新了工业机器人作为机器、在系统集成前适用的安全要求。"},"implication":{"en":"Robot manufacturers should align risk reduction, design evidence and user information early, not at final certification.","zh":"机器人制造商应尽早对齐风险降低、设计证据和用户信息，而不是等到最终认证。"}},{"id":"monitoring-will-be-risk-tiered","title":{"en":"Monitoring will become risk-tiered","zh":"监测将按风险分层"},"source":"NIST CAISI","sourceTitle":"Challenges to the Monitoring of Deployed AI Systems: Center for AI Standards and Innovation","sourceUrl":"https://www.nist.gov/publications/challenges-monitoring-deployed-ai-systems-center-ai-standards-and-innovation","published":"2026-03-06","reviewed":"2026-07-26","status":"forecast","horizon":"1-3y","relation":"supports-or-limits","summary":{"en":"NIST highlights unresolved questions about cadence, responsibility and risk level in deployed-system monitoring.","zh":"NIST指出部署后监测仍需解决频率、责任和风险等级等问题。"},"implication":{"en":"Routine drafting agents may be sampled, while financial, medical or physical actions may require near-continuous supervision.","zh":"普通写作Agent可以抽样检查，而金融、医疗或物理行动可能需要近乎持续监督。"}}],"relatedAnalysis":[{"id":"ai-in-robotics","path":"/news/insights/ai-in-robotics","title":{"en":"From Perception to Autonomy: Where AI Is Changing Robotics","zh":"从感知到自主：AI 正在如何改变机器人"},"sourceUrl":"https://ifr.org/ifr-press-releases/news/position-paper-on-ai-in-robotics","reviewed":"27 July 2026"},{"id":"international-ai-safety-report-2026","path":"/news/insights/international-ai-safety-report-2026","title":{"en":"Capability Is Advancing Faster Than Our Ability to Measure Every Risk","zh":"AI 能力的进步，正在快于我们衡量全部风险的能力"},"sourceUrl":"https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026","reviewed":"27 July 2026"},{"id":"responsible-agentic-ai","path":"/news/insights/responsible-agentic-ai","title":{"en":"From Singapore: Foundations for Responsible Agentic AI","zh":"从新加坡出发：负责任部署 Agentic AI 的基础"},"sourceUrl":"https://www.imda.gov.sg/resources/press-releases-factsheets-and-speeches/factsheets/2026/updated-model-ai-governance-framework-for-agentic-ai","reviewed":"27 July 2026"}]},{"id":"ai-for-scientific-discovery-and-reproducibility","number":"04","version":"1.0","created":"2026-07-28","reviewed":"2026-07-28","field":{"en":"Scientific intelligence","zh":"科学智能"},"fieldPath":"/research/fields/causal-scientific-intelligence","question":{"en":"When does AI accelerate genuine scientific discovery, and what records are needed to make machine-assisted results reproducible?","zh":"AI 在什么条件下真正加速科学发现？机器辅助成果需要保留哪些记录才能被复核与复现？"},"shortTitle":{"en":"AI for scientific discovery and reproducibility","zh":"AI 科学发现与可复核性"},"evidenceStatus":{"id":"mixed","label":{"en":"Mixed","zh":"证据与解释存在分歧"},"explanation":{"en":"Credible evidence supports more than one interpretation, or outcomes vary materially by context.","zh":"可信证据支持不止一种解释，或结果会随场景、群体与实施方式显著变化。"}},"currentEvidencePosition":{"en":"AI is producing useful predictions, candidate algorithms, hypotheses and research assistance in domains with strong data or evaluators. The strongest evidence appears where outputs can be independently measured. General claims of autonomous discovery remain premature without prospective validation, provenance and reproducible experimental records.","zh":"在拥有高质量数据或可靠评估器的领域，AI 已能产生有用预测、候选算法、研究假设与科研辅助。最强证据来自结果可被独立测量的场景。若缺乏前瞻验证、来源记录与可复现实验档案，关于“自主发现”的广泛主张仍为时过早。"},"scope":{"en":"This dossier covers machine-assisted hypothesis generation, algorithm and biological discovery, evaluators, peer-review support, provenance, prospective validation and reproducibility. It distinguishes prediction from mechanism and candidate generation from confirmed discovery.","zh":"本档案覆盖机器辅助假设生成、算法与生物发现、评估器、同行评审辅助、来源证明、前瞻验证与可复核性；明确区分预测与机制、候选生成与已确认发现。"},"supportedClaims":[{"en":"AI can search large candidate spaces effectively when a fast, objective and relevant evaluator is available.","zh":"当存在快速、客观且相关的评估器时，AI 可以有效搜索大规模候选空间。"},{"en":"Models can narrow experimental search, predict structure or regulation and support literature synthesis.","zh":"模型可以缩小实验搜索空间、预测结构或调控关系，并辅助文献综合。"},{"en":"Automated checks can strengthen parts of research review, especially formal, statistical and consistency checks.","zh":"自动化检查可以加强科研审阅的部分环节，特别是形式、统计与一致性检查。"},{"en":"Provenance and disclosure of model contribution are necessary for scientific assessment and attribution.","zh":"模型贡献的来源记录与透明披露，是科学评估和成果归属所必需的。"}],"evidenceTensions":[{"en":"A benchmark or evaluator may be easy to optimise without capturing the scientific property that ultimately matters.","zh":"基准或评估器可能容易被优化，却未必真正衡量最终重要的科学属性。"},{"en":"Prediction can guide experiments without establishing mechanism, causality, safety or clinical validity.","zh":"预测可以指导实验，但不能自动证明机制、因果关系、安全性或临床有效性。"},{"en":"AI-assisted review can catch errors while introducing automation bias, homogenised judgment or undisclosed model dependence.","zh":"AI 辅助审阅可以发现错误，也可能带来自动化偏见、判断同质化或未披露的模型依赖。"}],"openQuestions":[{"en":"Which domains have evaluators strong enough for closed-loop discovery, and which still require costly physical validation?","zh":"哪些领域的评估器足以支持闭环发现？哪些领域仍必须依赖高成本的物理验证？"},{"en":"How should journals, laboratories and funders record model, prompt, tool and human contributions?","zh":"期刊、实验室与资助机构应如何记录模型、提示、工具与人类的贡献？"},{"en":"Can AI-generated hypotheses produce durable novelty rather than recombinations favoured by existing literature?","zh":"AI 生成的假设能否形成持久的新颖性，而不只是重组现有文献偏好的方向？"}],"testableQuestions":[{"en":"Can an independent team reproduce the result from the published data, evaluation procedure and contribution record?","zh":"独立团队能否依据公开数据、评估程序与贡献记录复现结果？"},{"en":"Does prospective testing confirm performance on data and conditions unavailable during model development?","zh":"前瞻测试能否在模型开发期间不可见的数据与条件中确认性能？"},{"en":"Does AI assistance improve validated discoveries per unit of time or cost, rather than only generating more candidates?","zh":"AI 辅助能否提高单位时间或成本内经验证的发现数量，而不只是产生更多候选？"}],"decisionRelevance":[{"en":"Research: publish evaluator design, negative results, rejected candidates and the full AI contribution chain.","zh":"科研：公开评估器设计、负面结果、被淘汰候选与完整 AI 贡献链。"},{"en":"Institutions: invest in shared test beds, independent replication and durable research records—not models alone.","zh":"机构：投资共享测试平台、独立复现与长期科研档案，而不只投资模型。"},{"en":"Communication: describe predictions and candidates precisely; do not promote them as validated mechanisms or treatments.","zh":"传播：准确说明预测与候选结果，不把它们宣传为已验证机制或治疗方案。"}],"primaryEvidence":[{"id":"verified-search-loops-move-from-demo-to-discovery","title":{"en":"Verified search loops are becoming a practical discovery method","zh":"“生成—验证—进化”正在成为可落地的发现方法"},"source":"Google DeepMind","sourceTitle":"AlphaEvolve: How our Gemini-powered coding agent is scaling impact across fields","sourceUrl":"https://deepmind.google/blog/alphaevolve-impact/","published":"2026-05-07","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"AlphaEvolve shows a repeatable pattern: models propose candidate programs, objective evaluators test them, and an evolutionary loop retains better solutions. Its reported applications now span computing, mathematics, genomics, power systems and Earth science.","zh":"AlphaEvolve 展示了一种可重复的路径：模型提出候选程序，客观评估器进行测试，进化循环保留更优解。公开案例已从计算与数学延伸至基因组、电力系统和地球科学。"},"implication":{"en":"The frontier is shifting from AI that merely suggests ideas to systems that can search large solution spaces when results are automatically measurable.","zh":"前沿正在从“AI 提供想法”转向“AI 在结果可自动测量的领域搜索巨大解空间”。"}},{"id":"evaluators-become-core-discovery-infrastructure","title":{"en":"Evaluators may matter as much as generators","zh":"评估器的重要性可能不亚于生成模型"},"source":"Google DeepMind","sourceTitle":"AlphaEvolve: How our Gemini-powered coding agent is scaling impact across fields","sourceUrl":"https://deepmind.google/blog/alphaevolve-impact/","published":"2026-05-07","reviewed":"2026-07-26","status":"emerging","horizon":"1-3y","relation":"supports-or-limits","summary":{"en":"The value of an algorithm-discovery agent depends on whether candidate outputs can be tested quickly, consistently and at scale. Better evaluators can turn broad model creativity into dependable experimental progress.","zh":"算法发现 Agent 的价值取决于候选结果能否被快速、一致并规模化地测试。更好的评估器可以把模型的广泛创造力转化为更可靠的实验进展。"},"implication":{"en":"Research platforms will increasingly compete on verification environments, datasets and scoring systems—not only on the model that generates candidates.","zh":"未来研究平台的竞争重点将不仅是生成模型，也包括验证环境、数据集和评分系统。"}},{"id":"machine-discovered-methods-require-provenance","title":{"en":"Machine-discovered methods will need auditable provenance","zh":"机器发现的方法将需要可审计的来源链"},"source":"Google DeepMind","sourceTitle":"AlphaEvolve: How our Gemini-powered coding agent is scaling impact across fields","sourceUrl":"https://deepmind.google/blog/alphaevolve-impact/","published":"2026-05-07","reviewed":"2026-07-26","status":"forecast","horizon":"3-7y","relation":"supports-or-limits","summary":{"en":"As autonomous search contributes to consequential engineering and scientific results, knowing which model, prompt, evaluator, data and human decision produced a method becomes part of its credibility.","zh":"当自主搜索参与重要工程与科研成果时，所用模型、提示、评估器、数据及人类决策将成为成果可信度的一部分。"},"implication":{"en":"Within several years, research-grade discovery systems are likely to treat provenance, reproducibility and rollback records as default infrastructure.","zh":"未来数年，研究级发现系统很可能把来源追踪、可复现性和回滚记录作为默认基础设施。"}},{"id":"ai-pre-review-becomes-research-quality-layer","title":{"en":"AI pre-review is emerging as a research quality layer","zh":"AI 预审正在成为科研质量的新层级"},"source":"Stanford HAI","sourceTitle":"AI’s Growing Role as Scientific Peer Reviewer","sourceUrl":"https://hai.stanford.edu/news/ais-growing-role-as-scientific-peer-reviewer","published":"2026-03-25","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"Stanford reports large-scale experiments in which AI assistance supported scientific review. The strongest use today is early feedback on gaps, inconsistencies and technical issues before formal submission.","zh":"斯坦福介绍了在大规模科学评审中使用 AI 辅助的实验。当前最成熟的用途，是在正式投稿前发现缺口、不一致和技术问题。"},"implication":{"en":"Researchers can treat AI as a first-pass critic while keeping scientific claims and final editorial decisions accountable to people.","zh":"科研人员可以把 AI 作为第一轮批评者，但科学主张与最终编辑决定仍由人承担责任。"}},{"id":"scientific-ai-needs-prospective-validation","title":{"en":"Scientific AI needs prospective validation","zh":"科学 AI 需要前瞻性验证"},"source":"Google Research","sourceTitle":"Accelerating scientific breakthroughs with an AI co-scientist","sourceUrl":"https://research.google/blog/accelerating-scientific-breakthroughs-with-an-ai-co-scientist/","published":"2025-02-19","reviewed":"2026-07-26","status":"emerging","horizon":"1-3y","relation":"supports-or-limits","summary":{"en":"Retrospective rediscovery can show that a system recognizes useful patterns, but the stronger test is whether a new hypothesis survives pre-registered experiments and independent replication.","zh":"回顾性重新发现能够说明系统识别了有用模式，但更强的检验是新假设能否经受预注册实验和独立复现。"},"implication":{"en":"Credible co-scientist platforms should separate generated proposals, expert selection, experimental tests and confirmed findings in their public record.","zh":"可信的共同科学家平台应在公开记录中区分生成建议、专家选择、实验检验与已经确认的发现。"}},{"id":"genomic-ai-remains-a-hypothesis-engine","title":{"en":"Genomic AI remains a hypothesis engine, not an oracle","zh":"基因组 AI 仍是提出假设的引擎，而不是神谕"},"source":"Nature","sourceTitle":"Advancing regulatory variant effect prediction with AlphaGenome","sourceUrl":"https://www.nature.com/articles/s41586-025-10014-0","published":"2026-01-28","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"The paper reports broad predictive capability while also documenting limits and benchmark-dependent performance. Biological systems contain context and causal interactions that computational scores may miss.","zh":"论文报告了广泛的预测能力，同时也说明了局限及依赖基准的表现。生物系统包含模型评分可能遗漏的情境与因果交互。"},"implication":{"en":"Responsible use should present outputs as ranked hypotheses, with uncertainty and domain limits visible to researchers and clinicians.","zh":"负责任的使用方式应把输出呈现为带排序的假设，并向科研人员和临床人员清楚展示不确定性与适用边界。"}}],"relatedAnalysis":[{"id":"ai-for-scientific-discovery","path":"/news/insights/ai-for-scientific-discovery","title":{"en":"AI Is Becoming an Instrument of Scientific Discovery","zh":"人工智能正在成为科学发现的新型工具"},"sourceUrl":"https://www.nature.com/articles/s41586-024-07487-w","reviewed":"27 July 2026"},{"id":"international-ai-safety-report-2026","path":"/news/insights/international-ai-safety-report-2026","title":{"en":"Capability Is Advancing Faster Than Our Ability to Measure Every Risk","zh":"AI 能力的进步，正在快于我们衡量全部风险的能力"},"sourceUrl":"https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026","reviewed":"27 July 2026"},{"id":"multimodal-ai-cultural-diversity","path":"/news/insights/multimodal-ai-cultural-diversity","title":{"en":"Global Multimodal AI Must Look Beyond Western Benchmarks","zh":"面向全球的多模态 AI，不能只优化西方基准"},"sourceUrl":"https://deepmind.google/research/publications/132991/","reviewed":"27 July 2026"}]},{"id":"ai-infrastructure-energy-and-compute","number":"05","version":"1.0","created":"2026-07-28","reviewed":"2026-07-28","field":{"en":"Physical infrastructure","zh":"物理基础设施"},"fieldPath":"/research/fields/ai-economy","question":{"en":"How can AI capability grow without making energy, grid capacity, water, chips and geographic concentration invisible externalities?","zh":"AI 能力应如何增长，才能避免把能源、电网容量、水资源、芯片与地域集中度变成被忽视的外部成本？"},"shortTitle":{"en":"AI infrastructure, energy and compute","zh":"AI 基础设施、能源与算力"},"evidenceStatus":{"id":"developing","label":{"en":"Developing","zh":"证据正在积累"},"explanation":{"en":"The direction is visible, but the evidence base, operating history or independent validation remains incomplete.","zh":"方向已经可见，但证据基础、运行历史或独立验证仍不完整。"}},"currentEvidencePosition":{"en":"AI growth is now inseparable from electricity systems, data-centre geography, cooling, chips, networks and capital allocation. Efficiency and flexible compute can reduce some constraints, yet rebound effects, local grid bottlenecks and supply concentration make the net outcome uncertain.","zh":"AI 增长已经无法与电力系统、数据中心地理位置、冷却、芯片、网络和资本配置分开讨论。效率提升与灵活算力可以缓解部分约束，但反弹效应、地方电网瓶颈与供应集中使净影响仍不确定。"},"scope":{"en":"This dossier examines electricity demand, grid integration, data-centre clustering, efficiency, flexible workloads, chips, networks and supply-chain concentration. It does not publish a single universal footprint number because impacts vary by model, location, time and energy mix.","zh":"本档案研究用电需求、电网接入、数据中心集聚、效率、灵活负载、芯片、网络与供应链集中；不会给出单一“通用足迹”数字，因为影响会随模型、地点、时间与能源结构变化。"},"supportedClaims":[{"en":"Data-centre electricity demand is growing and increasingly material to grid planning in several regions.","zh":"数据中心用电需求正在增长，并已成为多个地区电网规划的重要因素。"},{"en":"Local concentration matters: a manageable global share can still create severe regional capacity and connection constraints.","zh":"地域集中很重要：全球占比看似可控，也可能在局部造成严重容量与接入约束。"},{"en":"Efficiency improvements can reduce energy per task, while lower cost and higher demand may increase total use.","zh":"效率提升可以降低单位任务能耗，但成本下降与需求增加也可能推高总用量。"},{"en":"Schedulable workloads and co-optimisation can make compute a more flexible grid participant in suitable settings.","zh":"在适当场景下，可调度负载与协同优化可以让算力成为更灵活的电网参与者。"}],"evidenceTensions":[{"en":"Faster chips and models improve efficiency while accelerating the scale and frequency of AI use.","zh":"更高效的芯片与模型可以节能，也可能加速 AI 使用规模和频率增长。"},{"en":"Low-carbon power procurement can support clean generation, but contractual claims may not match hourly local system impact.","zh":"采购低碳电力可以支持清洁能源，但合同层面的声明未必等同于当地电力系统每小时的真实影响。"},{"en":"Large clusters can improve economics and research capacity while concentrating supply, security and community risk.","zh":"大型集群可以改善经济性与科研能力，也会集中供应、安全与社区风险。"}],"openQuestions":[{"en":"How quickly will inference demand, agentic workloads and embodied systems change the balance between training and serving?","zh":"推理需求、Agent 工作负载与具身系统将以多快速度改变训练与服务之间的资源比例？"},{"en":"Which efficiency gains will reduce total resource use rather than being absorbed by higher demand?","zh":"哪些效率提升会真正减少总资源使用，而不是被更高需求抵消？"},{"en":"How should infrastructure disclosures represent location, time, water, hardware lifecycle and avoided impact together?","zh":"基础设施披露应如何同时呈现地点、时间、水资源、硬件生命周期与被避免的影响？"}],"testableQuestions":[{"en":"Can the operator report energy, carbon, water and hardware use at a decision-relevant temporal and geographic resolution?","zh":"运营方能否按对决策有意义的时间与地域粒度报告能源、碳、水和硬件使用？"},{"en":"Can flexible workloads reduce peak strain without degrading reliability, security or research validity?","zh":"灵活负载能否降低峰值压力，同时不损害可靠性、安全或科研有效性？"},{"en":"Does a claimed efficiency improvement reduce lifecycle impact after demand growth and infrastructure expansion are included?","zh":"在计入需求增长与基础设施扩张后，所声称的效率提升是否真正降低全生命周期影响？"}],"decisionRelevance":[{"en":"Strategy: treat energy, grid connection, chips, networking and location as core AI architecture decisions.","zh":"战略：把能源、电网接入、芯片、网络与地点视为 AI 架构的核心决策。"},{"en":"Engineering: optimise the complete workload and infrastructure system, not a model benchmark in isolation.","zh":"工程：优化完整工作负载与基础设施系统，而不是孤立的模型基准。"},{"en":"Disclosure: report boundaries, geography, time basis and uncertainty so footprint claims can be interpreted.","zh":"披露：说明边界、地域、时间基准与不确定性，使资源足迹声明可被正确理解。"}],"primaryEvidence":[{"id":"ai-power-demand-is-now-an-operating-constraint","title":{"en":"AI power demand is now an operating constraint","zh":"电力需求已成为AI运营约束"},"source":"International Energy Agency","sourceTitle":"Key Questions on Energy and AI","sourceUrl":"https://www.iea.org/reports/key-questions-on-energy-and-ai","published":"2026-04-16","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"The IEA’s 2026 update treats electricity availability, grid timing and local concentration as central variables in AI expansion. Compute plans can no longer assume power arrives automatically.","zh":"IEA的2026年更新把电力可用性、电网时序和区域集中度视为AI扩张的核心变量。算力计划不能再假定电力会自动到位。"},"implication":{"en":"Capacity strategy should model location, interconnection and energy delivery before hardware procurement.","zh":"采购硬件前，应先建模选址、并网和能源交付能力。"}},{"id":"grid-timelines-may-outlast-chip-cycles","title":{"en":"Grid timelines may outlast chip cycles","zh":"电网建设周期可能长于芯片迭代周期"},"source":"International Energy Agency","sourceTitle":"Key Questions on Energy and AI","sourceUrl":"https://www.iea.org/reports/key-questions-on-energy-and-ai","published":"2026-04-16","reviewed":"2026-07-26","status":"emerging","horizon":"3-7y","relation":"supports-or-limits","summary":{"en":"AI hardware advances quickly, while transmission, generation and interconnection projects often take years. This mismatch can strand equipment plans or concentrate development in power-ready regions.","zh":"AI硬件进步很快，而输电、发电和并网项目通常需要多年。这种错配可能让设备计划搁浅，或使开发集中于电力条件成熟的地区。"},"implication":{"en":"Long-term energy partnerships may be as strategic as short-term access to accelerators.","zh":"长期能源合作可能与短期获取加速芯片同样具有战略意义。"}},{"id":"data-centre-clusters-create-local-system-risks","title":{"en":"Data-centre clusters create local system risks","zh":"数据中心集群会形成地方系统风险"},"source":"International Energy Agency","sourceTitle":"Key Questions on Energy and AI","sourceUrl":"https://www.iea.org/reports/key-questions-on-energy-and-ai","published":"2026-04-16","reviewed":"2026-07-26","status":"emerging","horizon":"1-3y","relation":"supports-or-limits","summary":{"en":"Global electricity totals can hide severe local pressure where facilities cluster. Network congestion, generation mix, water availability and community acceptance can determine project viability.","zh":"全球电力总量可能掩盖设施聚集地区的严重压力。电网拥堵、能源结构、水资源和社区接受度都可能决定项目可行性。"},"implication":{"en":"Infrastructure disclosure should be regional and site-specific, not only global.","zh":"基础设施披露应具体到区域和站点，而不只是全球总量。"}},{"id":"efficiency-can-increase-total-ai-use","title":{"en":"Efficiency can increase total AI use","zh":"效率提升可能反而增加AI总使用量"},"source":"International Energy Agency","sourceTitle":"Key Questions on Energy and AI","sourceUrl":"https://www.iea.org/reports/key-questions-on-energy-and-ai","published":"2026-04-16","reviewed":"2026-07-26","status":"forecast","horizon":"3-7y","relation":"supports-or-limits","summary":{"en":"Cheaper and more efficient computation reduces the energy needed for each task but can also stimulate far more usage. System demand may rise even while unit efficiency improves.","zh":"更便宜、更高效的计算会降低单项任务所需能源，但也可能刺激更大规模的使用。即使单位效率改善，系统总需求仍可能上升。"},"implication":{"en":"Energy planning should model rebound effects rather than extrapolate per-query savings alone.","zh":"能源规划应考虑反弹效应，而不能只外推单次查询的节省。"}},{"id":"flexible-compute-can-become-a-grid-resource","title":{"en":"Flexible compute can become a grid resource","zh":"可调度算力可以成为电网资源"},"source":"International Energy Agency","sourceTitle":"Key Questions on Energy and AI","sourceUrl":"https://www.iea.org/reports/key-questions-on-energy-and-ai","published":"2026-04-16","reviewed":"2026-07-26","status":"forecast","horizon":"3-7y","relation":"supports-or-limits","summary":{"en":"Some training and batch workloads can move across time or location more easily than conventional industrial demand. With safeguards, scheduling flexibility could help align compute with cleaner and less constrained electricity.","zh":"部分训练和批处理负载比传统工业需求更容易在时间或地点上移动。在具备保障措施时，调度灵活性可使算力更好匹配清洁且不拥堵的电力。"},"implication":{"en":"Future cloud contracts may value when and where computation runs, not only speed and price.","zh":"未来云合同可能不仅关心速度与价格，也会重视计算在何时何地运行。"}},{"id":"ai-and-energy-are-a-two-way-system","title":{"en":"AI and energy form a two-way system","zh":"AI与能源构成双向系统"},"source":"International Energy Agency","sourceTitle":"Energy and AI","sourceUrl":"https://www.iea.org/reports/energy-and-ai","published":"2025-04-10","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"The IEA analyses both electricity required by AI and the use of AI to improve energy operations. Costs and benefits therefore belong in the same system view rather than separate debates.","zh":"IEA同时分析AI所需电力以及AI改善能源运营的用途。因此，成本与收益应被放在同一个系统视角中，而不是分开讨论。"},"implication":{"en":"Responsible strategy should count consumption, resilience and optimisation together.","zh":"负责任的战略应同时计算消耗、韧性和优化收益。"}}],"relatedAnalysis":[{"id":"energy-and-ai-infrastructure","path":"/news/insights/energy-and-ai-infrastructure","title":{"en":"AI Civilization Has a Physical Energy System Beneath It","zh":"AI 文明的底层，是一套真实存在的能源系统"},"sourceUrl":"https://www.iea.org/reports/energy-and-ai/","reviewed":"27 July 2026"},{"id":"ai-supply-chain","path":"/news/insights/ai-supply-chain","title":{"en":"The AI Economy Is Also a Five-Layer Supply Chain","zh":"AI 经济也是一条由五个层级组成的供应链"},"sourceUrl":"https://www.bis.org/publ/bppdf/bispap154.htm","reviewed":"27 July 2026"},{"id":"global-ai-index-2026","path":"/news/insights/global-ai-index-2026","title":{"en":"AI Adoption Is Accelerating, While Agent Deployment Remains Early","zh":"AI 采用正在加速，但 Agent 部署仍处早期"},"sourceUrl":"https://hai.stanford.edu/ai-index/2026-ai-index-report","reviewed":"27 July 2026"}]},{"id":"human-ai-collaboration-and-decision-quality","number":"06","version":"1.0","created":"2026-07-28","reviewed":"2026-07-28","field":{"en":"Human–AI collaboration","zh":"人机协作"},"fieldPath":"/research/fields/human-ai-collaboration","question":{"en":"Which forms of human–AI collaboration improve decision quality, capability and agency—and for whom?","zh":"哪些人机协作方式能够提升决策质量、能力与人的主体性？这些收益又真正属于谁？"},"shortTitle":{"en":"Human–AI collaboration and decision quality","zh":"人机协作与决策质量"},"evidenceStatus":{"id":"mixed","label":{"en":"Mixed","zh":"证据与解释存在分歧"},"explanation":{"en":"Credible evidence supports more than one interpretation, or outcomes vary materially by context.","zh":"可信证据支持不止一种解释，或结果会随场景、群体与实施方式显著变化。"}},"currentEvidencePosition":{"en":"Field evidence shows meaningful gains in some tasks, including the diffusion of expertise, while effects remain uneven across workers, workflows and decision types. Tool access alone does not redesign work. Outcomes depend on task structure, user skill, organisational change, feedback and accountable human judgment.","zh":"现实证据显示，AI 在部分任务中能够带来显著收益，包括扩散专业经验；但不同人员、工作流与决策类型的效果并不均衡。仅仅获得工具并不会自动重塑工作，结果取决于任务结构、使用者技能、组织变革、反馈机制与可问责的人类判断。"},"scope":{"en":"This dossier examines field studies, labour evidence, expertise diffusion, workflow redesign, oversight, skill development and human agency. It does not generalise results from one occupation, organisation or model to all work.","zh":"本档案研究现场实验、劳动证据、专业经验扩散、工作流重构、监督、技能发展与人的主体性；不会把单一职业、组织或模型的结果推广到所有工作。"},"supportedClaims":[{"en":"AI assistance can improve productivity and quality in bounded tasks, with some studies showing larger gains for less-experienced workers.","zh":"AI 辅助可以在边界明确的任务中提升效率与质量，部分研究显示经验较少的员工可能获得更大收益。"},{"en":"Exposure is more likely to transform many jobs than eliminate every task within them.","zh":"对许多工作而言，AI 更可能改变任务组合，而不是消除其中全部任务。"},{"en":"Workflow and organisational redesign are necessary to convert tool use into durable performance improvement.","zh":"要把工具使用转化为持久绩效提升，需要同步重构工作流与组织方式。"},{"en":"Human accountability remains necessary for consequential decisions even when AI performs substantial analysis.","zh":"即使 AI 承担大量分析，重要决策仍需要由人类承担最终责任。"}],"evidenceTensions":[{"en":"Assistance can spread expertise while also encouraging overreliance, skill atrophy or convergence on model-preferred answers.","zh":"辅助工具可以扩散专业能力，也可能导致过度依赖、技能退化或判断趋同于模型偏好的答案。"},{"en":"Average productivity gains can conceal unequal effects by experience, gender, income, language or access.","zh":"平均效率收益可能掩盖经验、性别、收入、语言与可获得性方面的不平等影响。"},{"en":"Human review can be a real control or a ceremonial step, depending on time, competence, information and authority.","zh":"人类复核可能是真实控制，也可能只是形式步骤，取决于时间、能力、信息与权限。"}],"openQuestions":[{"en":"Which collaboration patterns create durable learning rather than short-term output gains?","zh":"哪些协作模式能够形成持久学习，而不只是短期产出提升？"},{"en":"How do repeated AI-mediated decisions change professional judgment, confidence and institutional memory?","zh":"反复由 AI 介入的决策将如何改变专业判断、自信与组织记忆？"},{"en":"Which groups bear transition costs, and which governance mechanisms distribute gains more fairly?","zh":"哪些群体承担转型成本？哪些治理机制能更公平地分配收益？"}],"testableQuestions":[{"en":"Does the human–AI team outperform both the unaided human and the AI alone on quality, not only speed?","zh":"在人类独立工作和 AI 独立工作之外，人机团队能否在质量而不仅是速度上表现更好？"},{"en":"Can users detect model error, explain their final decision and disagree without procedural penalty?","zh":"使用者能否识别模型错误、解释最终决定，并在不同意模型时不受到流程性惩罚？"},{"en":"Do gains remain after novelty effects, selection bias and implementation support are removed?","zh":"排除新奇效应、选择偏差与实施支持后，收益是否仍然存在？"}],"decisionRelevance":[{"en":"Research: measure decision quality, learning, equity and agency alongside output and speed.","zh":"科研：在产出与速度之外，同时衡量决策质量、学习、公平与主体性。"},{"en":"Organisations: redesign roles, feedback, escalation and training rather than layering AI onto an unchanged process.","zh":"组织：重构角色、反馈、升级机制与培训，而不是把 AI 叠加在不变的流程上。"},{"en":"Governance: ensure affected people can understand, contest and obtain human review of consequential decisions.","zh":"治理：确保受影响者能够理解、质疑重要决策，并获得真正的人类复核。"}],"primaryEvidence":[{"id":"augmentation-may-compress-some-skill-gaps","title":{"en":"Augmentation may compress some skill gaps","zh":"增强型AI可能缩小部分技能差距"},"source":"The Quarterly Journal of Economics","sourceTitle":"Generative AI at Work","sourceUrl":"https://academic.oup.com/qje/article/140/2/889/7990658","published":"2025-02-04","reviewed":"2026-07-26","status":"forecast","horizon":"3-7y","relation":"supports-or-limits","summary":{"en":"FUURAA forecasts that well-designed assistance will help more people reach competent performance in bounded tasks, while shifting the frontier toward judgement and responsibility.","zh":"FUURAA预测，设计良好的辅助将在有限任务中帮助更多人达到合格水平，同时把专家前沿推向判断与责任。"},"implication":{"en":"The value of expertise may move from producing routine answers to handling exceptions and setting standards.","zh":"专业价值可能从生产常规答案转向处理例外和制定标准。"}},{"id":"individual-tools-do-not-redesign-jobs-alone","title":{"en":"Individual tools do not redesign jobs alone","zh":"个人工具不会自动重构岗位"},"source":"National Bureau of Economic Research","sourceTitle":"Shifting Work Patterns with Generative AI","sourceUrl":"https://www.nber.org/papers/w33795","published":"2025-05-09","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"The experiment detected personal time savings without clear shifts in the quantity or composition of tasks. Giving individuals a tool is different from redesigning a team, process or service.","zh":"实验发现个人节省了时间，但任务数量或构成没有明显变化。给个人一个工具，与重构团队、流程或服务是两件不同的事。"},"implication":{"en":"Enterprise transformation requires coordinated workflow decisions in addition to employee access.","zh":"企业转型除了员工获得工具，还需要协调一致的流程决策。"}},{"id":"job-transformation-is-more-likely-than-full-replacement","title":{"en":"Job transformation is more likely than full replacement","zh":"岗位转型比整体替代更可能发生"},"source":"International Labour Organization","sourceTitle":"Generative AI and Jobs: A Refined Global Index of Occupational Exposure","sourceUrl":"https://www.ilo.org/publications/generative-ai-and-jobs-refined-global-index-occupational-exposure","published":"2025-05-20","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"Most exposed occupations contain a mixture of tasks, so current evidence points more strongly to changing task bundles than removing entire occupations.","zh":"大多数暴露职业都包含多种任务，因此现有证据更支持任务组合变化，而不是整个职业消失。"},"implication":{"en":"Training and workflow redesign may matter sooner than forecasts of whole-job automation.","zh":"培训和工作流重构的重要性，可能早于对整岗自动化的预测。"}},{"id":"task-redesign-and-social-dialogue-will-shape-outcomes","title":{"en":"Task redesign and social dialogue will shape outcomes","zh":"任务重构与社会对话将决定转型结果"},"source":"International Labour Organization","sourceTitle":"Generative AI and Jobs: A Refined Global Index of Occupational Exposure","sourceUrl":"https://www.ilo.org/publications/generative-ai-and-jobs-refined-global-index-occupational-exposure","published":"2025-05-20","reviewed":"2026-07-26","status":"forecast","horizon":"3-7y","relation":"supports-or-limits","summary":{"en":"FUURAA forecasts that similar AI capability will produce different labour outcomes depending on worker voice, training access and how productivity gains are shared.","zh":"FUURAA预测，同样的AI能力会因员工参与、培训机会和生产力收益分配方式不同而产生不同劳动结果。"},"implication":{"en":"The future of work is partly an institutional design choice, not a technical result alone.","zh":"工作的未来部分取决于制度设计，而不只是技术结果。"}},{"id":"productivity-follows-the-jagged-frontier","title":{"en":"Productivity follows a jagged frontier","zh":"生产力提升沿着参差不齐的能力边界发生"},"source":"Stanford Institute for Human-Centered Artificial Intelligence","sourceTitle":"Economy | The 2026 AI Index Report","sourceUrl":"https://hai.stanford.edu/ai-index/2026-ai-index-report/economy","published":"2026-04-13","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"Evidence reviewed by Stanford shows larger gains in structured, measurable work and weaker performance where judgement is deeper or outputs are hard to verify. A single productivity claim cannot represent all occupations.","zh":"斯坦福汇总的证据显示，结构化、可衡量工作中的增益更大，而需要深入判断或难以核验的任务表现较弱。单一的生产力数字无法代表所有职业。"},"implication":{"en":"Deploy AI task by task and match review intensity to uncertainty and consequence.","zh":"应逐项任务部署AI，并按照不确定性与后果决定审核强度。"}},{"id":"human-accountability-remains-the-anchor","title":{"en":"Human accountability remains the anchor","zh":"人类问责仍是智能体治理的锚点"},"source":"Infocomm Media Development Authority of Singapore","sourceTitle":"Singapore Launches New Model AI Governance Framework for Agentic AI","sourceUrl":"https://www.imda.gov.sg/resources/press-releases-factsheets-and-speeches/press-releases/2026/new-model-ai-governance-framework-for-agentic-ai","published":"2026-01-22","reviewed":"2026-07-26","status":"observed","horizon":"now","relation":"supports-or-limits","summary":{"en":"Singapore's agentic AI framework places ultimate accountability with people and organisations even when software can plan and act autonomously.","zh":"新加坡的智能体AI治理框架强调，即使软件能够自主规划与行动，最终责任仍由人员和组织承担。"},"implication":{"en":"FUURAA should make ownership, escalation and final decision rights visible in every consequential agent workflow.","zh":"FUURAA应在每个重要智能体流程中明确责任人、升级路径与最终决定权。"}}],"relatedAnalysis":[{"id":"human-ai-collaboration-at-work","path":"/news/insights/human-ai-collaboration-at-work","title":{"en":"Real-World Evidence Shows AI Can Spread Expertise—Unevenly","zh":"真实工作研究显示：AI 可以传播经验，但收益并不平均"},"sourceUrl":"https://academic.oup.com/qje/article/140/2/889/7990658","reviewed":"27 July 2026"},{"id":"ai-and-future-of-work","path":"/news/insights/ai-and-future-of-work","title":{"en":"AI Is More Likely to Transform Work Than Simply Replace It","zh":"AI 更可能重塑工作，而不是简单取代工作"},"sourceUrl":"https://www.ilo.org/publications/generative-ai-and-jobs-refined-global-index-occupational-exposure","reviewed":"27 July 2026"},{"id":"skills-and-organisations","path":"/news/insights/skills-and-organisations","title":{"en":"The Future of Work Is a Redesign of Skills and Organisations","zh":"工作的未来，也是技能与组织的共同重构"},"sourceUrl":"https://www.weforum.org/publications/the-future-of-jobs-report-2025/digest/","reviewed":"27 July 2026"},{"id":"human-development-in-ai-era","path":"/news/insights/human-development-in-ai-era","title":{"en":"The Measure of AI Progress Is Whether It Expands Human Possibility","zh":"衡量 AI 进步，应看它是否扩大了人的可能性"},"sourceUrl":"https://www.undp.org/press-releases/human-development-progress-slows-35-year-low-according-un-development-programme-report","reviewed":"27 July 2026"}]}]}