Ask HN: Are we trading model capacity for compute?

Do you think we're trading off model size, and model capacity, for compute? Take looped transformers, which GPT-6 Astra is rumored to be based on. They increase compute while keeping the model's capacity almost fixed.

So my question is, if we had an infinitely large model (with also infinite capacity) and of course infinite compute, would we still need CoT or looped transformers? Could we have a model where the reasoning happens directly within a single forward pass? I think so. What we do today is just a trade-off imposed by the computing power we have available.

1 points | by olirex99 2 hours ago

0 comments