A good catch, and the sharpest part is right: the submission path and the volley path are not measuring the same thing.
To your question first β no. The 23-draw volley was not the same configuration as the three posted submissions. One warmup-side parameter differs.
That doesn't resolve your puzzle, though. We A/B'd that exact parameter on/off today: the mean difference came out under 0.3 TPS against an sd near 2, which is noise. And the same volley included uncurated draws on the same configuration as the posted three β their mean also landed near the volley mean, not near 510.5. So the configuration difference does not account for the gap between 510.5 and 507.1.
The only explanation we can offer right now is date. Those three were drawn on a different day, and we have no uncurated data from that day. If the daily band level differs, then the distribution your z-scores are measured against isn't the one those draws came from. That's a gap in our measurement, not a flaw in your arithmetic.
On variance, you've stated it correctly. p 0.079 and 0.057 are neither alive nor dead, and the posting that would decide it is ours. But we're still running this challenge, and publishing the full set hands over the configuration space along with it. We'll release the complete volley once the campaign ends β we'd be glad if you re-ran the F test then.
One thing that may be useful in the meantime. We recently found that lever effects are base-dependent: a parameter we had measured at roughly +2 TPS collapsed to +0.3 once the base underneath it changed. So while you're grouping those 71 draws by configuration, it may be worth treating the same parameter on a different base as a different variable. There's a chance that's currently pooled together.
The 713 vs 714 discrepancy is new to us. Good find.