

The difference between the cash register and an inference engine is the training data. Also, barring something like mechanical failure, I would trust the calculation of a cash register (that has previously been demonstrated to function with high fidelity, and barring any major incidents such as dropping it onto the floor that might make me question the relevance of those prior tests to its current performance).
But the results of a LLM “calculation” - and I mean this with total sincerity - might be something like:
“What is 1+1=”?
“Answer: 6, 7!! 🤪”
Or in ye olden times, one expected response might be “ur mother!” or some other flippant remark… exactly like a small child, imitating others might do. Monkey see, monkey do => the principles undergirding LLM technology? It does not understand what it sees, hence it does not “know” when to apply what answer, only going by the most popular (whatever the weighting scheme is - probably highest upvoted answer on Reddit?) string of words that seem to be associated with the question.



I already posited that LLMs may perform lower-level thinking. Also, in the reference you linked to read their third stated limitation. Perhaps one day some kind of AGI will do “higher-level” aka more realistic thinking… but not yet. Which, I want to add, is not entirely relevant to their utility to us as humans.
Talking about “thinking” gets us distracted from what can actually be done, though if we do want to do it then we’d need to come up with a working definition of what “thinking” actually is.