Real-vs-noise verdicts
AI answers change from run to run, so a single score is one random draw. Orem samples each prompt many times, bootstraps a 95% confidence interval, and tells you plainly whether this week's change is real or within noise, so you never report or chase a move that was just randomness.