AI models that replace blood-flow simulation are scored by how well they match the flow field. That score does not tell you whether the residence-time biomarker doctors care about can be trusted. This page shows where it can't, on two real arteries, and what kind of error is to blame.
Blood at an artery wall pushes forward in systole and, in some places, backward in diastole. Two biomarkers summarise a heartbeat: TAWSS counts every step taken (10 + 9 = 19); RRT, the residence time, depends only on where you ended up (10 − 9 = 1).
Now let the model miscount one step. TAWSS is off by 1 in 19, about 5%. RRT is off by 1 in 1, 100%. Same small error, opposite consequences. Everything below is that arithmetic, measured on real simulations and a real AI model.
At each point of the wall, forward and backward shear partly cancel over a heartbeat. Time-averaged shear (TAWSS) adds them up, so errors stay proportionate. Residence time (RRT) depends on what survives the cancellation, so the same error is magnified. How much depends on the kind of error: the constants below were fitted for a clock/window error; a persistent spatial bias is punished about five times harder.
Same cerebral wall, same 10.8% field error, four ways of being wrong. Median excess RRT error over TAWSS on the most-reversing faces (cancellation factor above 1). A pure clock error integrated over the surrogate's own period changes nothing — a point owed to Wenhao Ding (Imperial College London), quoted below.
A POD + LSTM surrogate was trained on the cerebral CFD at six heart rates and asked to roll out the beat it had never seen. Its one-step accuracy is the kind of number a paper reports. Its rollout is what a clinician would actually use.
Leading flow mode, normalised. ■ CFD truth ■ surrogate rollout. The model reproduces the beat shape, then its clock pulls ahead by about 0.7 samples per beat.
The three numbers every surrogate paper already reports are a fingerprint. Perturb the same cerebral wall three different ways at a similar field error and score it the way the literature does. Then compare with an independent published model, whose fields we never touched.
Try your own model's numbers
The bands above were measured on a wall whose median OSI is 0.0007. They move with the geometry, so the fourth box matters: enter the median OSI of the wall your errors were computed on.
Typical values: intracranial aneurysm surfaces 0.0005–0.005; an abdominal aorta including non-reversing segments 0.0002–0.002; a sac or appendage taken alone 0.02–0.08. Above 0.05 this check is extrapolating past its calibration and says so.
Sent to twenty groups working on hemodynamic surrogates and biomarker robustness on 22 September 2026, with four questions. Replies are quoted with permission and the page is corrected where they were right.