bugfix — Equal-Weight Aggregation (companion to final-report.md)

Generated: 2026-05-19T04:05:40Z

Inputs and source artifacts

Same inputs as the canonical final-report.md — only the aggregation rule changes here.

Methodology

Ranking (Equal-Weight Mean)

  1. claudekit — 181.53/200
  2. ecc — 175.40/200
  3. pure — 172.03/200
  4. superpower — 169.43/200
  5. compound — 169.33/200
  6. bmad — 169.13/200
  7. omc — 167.97/200
  8. gstack — 164.69/200

Detail

Same cohort and judgments as final-report.md; only the aggregation rule differs. Column glossary:

Tool Equal-Weight Mean Pooled σ within_σ between_σ N
claudekit 181.53 14.33 11.42 9.35 75
ecc 175.40 17.20 13.54 11.70 75
pure 172.03 14.64 12.05 9.58 75
superpower 169.43 14.03 7.48 13.28 75
compound 169.33 13.79 9.57 11.27 75
bmad 169.13 17.15 12.75 12.94 75
omc 167.97 20.35 15.66 13.74 75
gstack 164.69 17.18 9.13 16.12 75

Cross-rule comparison

Compare Equal-Weight Mean here against Weighted Mean in final-report.md. Rank-1 is identical under both rules on every task in this corpus; mid-pack ranks 4–7 may swap by at most 2 positions.