refactor — Equal-Weight Aggregation (companion to final-report.md)

Generated: 2026-05-19T04:05:40Z

Inputs and source artifacts

Same inputs as the canonical final-report.md — only the aggregation rule changes here.

Methodology

Ranking (Equal-Weight Mean)

  1. pure — 182.63/200
  2. claudekit — 180.76/200
  3. superpower — 180.51/200
  4. bmad — 180.08/200
  5. compound — 177.03/200
  6. ecc — 176.57/200
  7. omc — 173.83/200
  8. gstack — 147.92/200

Detail

Same cohort and judgments as final-report.md; only the aggregation rule differs. Column glossary:

Tool Equal-Weight Mean Pooled σ within_σ between_σ N
pure 182.63 13.10 4.44 13.73 75
claudekit 180.76 16.25 5.29 17.01 75
superpower 180.51 14.74 5.09 15.39 75
bmad 180.08 14.86 5.48 15.28 75
compound 177.03 15.96 6.04 16.40 75
ecc 176.57 16.91 8.71 16.12 75
omc 173.83 17.50 7.43 17.52 75
gstack 147.92 58.70 58.43 12.52 75

Cross-rule comparison

Compare Equal-Weight Mean here against Weighted Mean in final-report.md. Rank-1 is identical under both rules on every task in this corpus; mid-pack ranks 4–7 may swap by at most 2 positions.