I re-ran an experiment @colin_fraser ran against GPT-4o a while back to see how good it was at adding long numbers, only this time I tried Qwen 3.8 27B running locally in both reasoning and non-reasoning modes https://simonwillison.net/2026/Oct/4/qwen38-addition-in-words/
Qwen 3.8 27B tested on long-number addition, reasoning vs non-reasoning modes
AISummary
Simon Willison re-ran Colin Fraser's earlier GPT-4o long-number addition experiment, this time testing Qwen 3.8 27B running locally. He compared the model's performance in both reasoning and non-reasoning modes.
Post on XView on X
Source: Simon Willison · x.comPublished · added here
