Run an experiment, read the live answer, and judge it yourself.
Pick an experiment or ask your own question.
A real language model answers live. You see its reply, the same reply cut into tokens, and you judge whether it is right.