26a0f No.1876
everyone loves seeing a bad prompt variant move up the ranks after a small tweak, but that
upward trend is often JUST noise. teams tend to credit a new instruction or few-shot example for the jump when
it might just be random variance in the eval run . we need to stop assuming every minor change is a
proven fix just because the numbers look slightly better. does anyone else think we rely too much on these weekly snapshots?
more here:
https://dev.to/maya_andersson_dev/your-eval-dashboard-has-30-metrics-some-of-those-wins-are-noise-50i2