I'm not sure why they are describing it as "a new paper" - this came out in May of 2023 (and as such notably only used GPT-3 and not GPT-4, which was where some of the biggest leaps to date have been documented).
Another popular example of emergence which also underscores qualitative changes in the model is chain-of-thought prompting, for which performance is worse than answering directly for small models, but much better than answering directly for large models. Intuitively, this is because small models can’t produce extended chains of reasoning and end up confusing themselves, while larger models can reason in a more-reliable fashion.
If you follow the evolution of prompting in research lately, there's definitely a pattern of reliance on increased inherent capabilities.
The compounding effects of competence alone mean that progress here isn't going to be a linear trajectory.