I don't think you're going to see 10 years of linear progress increasing the complexity of generated code, the same as you're not going to see a steady march towards factually correct LLM answers. We've had a step improvement towards producing certain types of text output, and people are racing to throw more resources at it for diminishing returns.
Producing code has got to be the best case scenario for LLMs because there's a lot of boilerplate and examples of simple applications, and people aren't out here criticizing the structure of a TODO list app.
There are often diminishing returns to naive scaling, but nobody knows whether we’re at the point of diminishing returns for machine learning research, broadly defined. Who knows what new tricks will be discovered? We don’t have any physical limits to guide us.
Producing code has got to be the best case scenario for LLMs because there's a lot of boilerplate and examples of simple applications, and people aren't out here criticizing the structure of a TODO list app.