Claim check True
Does more compute reliably make a better model? The scaling claim, tested
We plotted three labs' published benchmark scores against disclosed training compute. The line holds better than the memes suggest.
Draft entry. The full analysis for this piece ships with the data pipeline; the dataset it grades against is being assembled.