compare
Run every candidate over the same trace, so the comparison is genuinely like-for-like.
A fresh algorithm instance per candidate is the caller's responsibility — reusing one across runs is the classic way to get a result that cannot be reproduced.