I genuinely hate programming now

I know it’s not just pre-trained text in a chat interface. But either way it’s optimized for user desire because there is no notion of correctness that’s immediately available in the vast majority of real world applications.

That’s why it’s different with programming, because the result is immediately verifiable through just running the program or even through automated tests.

RLHF is also very faulty because it’s a yes/no for a whole generation. There’s no context or nuance you can give to teach it what specifically is right or wrong about a generation. You just nudge all weights in either direction.

You want anything else, you’re going to need human written pairs for fine-tuning. But as we know, that doesn’t scale.

RLVR is just not possible in any meaningful way in large models for a lot of context/project/domain-specific work. This is why I believe small models are the way forward.

The knowledge about and ease of training and fine-tuning small models needs to get a lot better before I see this changing.

As I mentioned, not comparable at all to free-form text generation imo. No matter the complexity, chess and go are closed-form problems and evaluating the correctness in a loop is as simple as looking at the rules.

Yes exactly. I want to clarify that we can still make advancements with AI through cross-context applications of knowledge. But that doesn’t mean we magically get AGI or superintelligence.