The same system can have faster first-token time, slower complete results, a worse latency tail and a lower completion rate. Evaluation must start with what the user waits for, pass through end-to-end clocks, production load and failure denominators, then return to a comparable, expiring conclusion.
Applicability boundaryThis is a public research method, not a FUURAA or FUUVO product-capability claim or an assessment of any model, provider, cloud platform, benchmark or monitoring tool. It is not an availability warranty, SLA, certification, audit, procurement, investment, legal or compliance opinion.