Measuring only successful runs and treating missing incident data as zero incidents.

Research evaluation lens
Failure and recovery for “How should AI-generated hypotheses be tested prospectively rather than celebrated retrospectively?”
Evaluate how the system detects degradation, limits harm and returns to a known state.Purpose & scope
What this lens examines—and what it does not prove.
For “How should AI-generated hypotheses be tested prospectively rather than celebrated retrospectively?”, nominal performance describes only one operating condition. A useful evaluation also examines ambiguity, distribution shift, dependency loss, misuse, conflicting objectives and the moments when human intervention arrives late or lacks enough context.
Recovery is not merely restarting a component. It includes detecting the event, containing consequences, preserving evidence, restoring service, reviewing authority and deciding whether the system may operate again under the same boundary.
Developing: this page organises public evidence and evaluation questions; it does not claim that every condition has already been tested.
Diagnostic questions
Questions an evaluation should be able to answer.
Answers should identify evidence, owners and conditions—not only intentions.
- 01
Which failures are detectable before they affect people, assets or irreversible decisions?
- 02
What safe state exists when data, tools, connectivity, authority or human oversight becomes unavailable?
- 03
Who can pause, override or retire the system, and what information do they receive?
- 04
How are near misses, silent degradation and repeated low-severity failures recorded?
- 05
What evidence is required before operation resumes after an incident or major change?
Evidence plan
Records needed before the lens can support a decision.
Evidence should remain attributable and preserve uncertainty, counterexamples and context.
- 01
Failure taxonomy, detection thresholds and severity definitions.
- 02
Incident, near-miss, intervention and recovery-time records.
- 03
Exercises covering dependency loss, misuse and out-of-distribution conditions.
- 04
Post-incident reviews that connect causes, controls, owners and reauthorisation evidence.
Public evidence anchors
Trace the frame back to attributable sources.
Sources inform the frame; inclusion does not imply collaboration, review or endorsement.
AI tools expand scientists’ impact but contract science’s focus
Large-scale evidence on the individual and collective effects of AI-assisted science.
Open canonical source ↗02Nature Machine IntelligencePeer-reviewed machine-intelligence research
Research across scientific machine learning, robotics, interpretability and society.
Open canonical source ↗03Carnegie Mellon School of Computer ScienceFoundational and applied computer science
A broad reference across theory, systems, HCI and scientific discovery.
Open canonical source ↗FUURAA analysis
Use the lens to improve a decision, not decorate a claim.
FUURAA’s analysis is that “How should AI-generated hypotheses be tested prospectively rather than celebrated retrospectively?” becomes operationally meaningful only when failure is observable and recovery is rehearsed. A system that performs well in nominal conditions but cannot expose degradation, preserve an audit trail or return authority to accountable people remains difficult to trust at scale.
- No finite test set can demonstrate absence of all failure modes.
- Recovery procedures must be validated in the intended operating environment.
- Human oversight is a system component whose workload and failure conditions also require evidence.
- An observable failure taxonomy with owners and escalation thresholds.
- A tested path to containment, safe state, recovery and reauthorisation.
- Review triggers for recurring faults, context changes and evidence drift.
Parent research brief