
There’s endless debate on X about which model has the best verbal judo.
LLMs aren’t magic.
They generate text by calculating probabilities across weighted values, learned patterns, and fragments of language. That alone is not impressive.
What matters is whether the model remains reliable when the problem becomes precise, technical, or deeply contextual.
Sometimes you don’t realize a model is terrible until it reaches a critical point and does something that destroys your trust.
By then, you’ve already paid for the service—and received nothing usable from it.
The problem is abstract, multidirectional, and almost comical in scope.
The practical solution is achievable:
1) Learn what works for you, what doesn’t, where the model fails, why it fails, and exactly which next step it cannot comprehend.
2) Open a new session with the same model and ask it to grade its own homework.
I hope nobody believes this—but if you do, hell yeah. Enjoy watching your $$ oscillate out the door. 😆
English
















