This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
Companies ship Italian that no Italian would write. In my own QA audits of software websites this year I found a transcription tool whose "Tools" menu said Utensileria (a hardware store), a homepage bragging about 5 Mld+ Altoparlanti nativi (five billion native loudspeakers), and more than one Fidato da, a word-for-word "Trusted by" that reads as machine output.
Models are now the reviewers many teams rely on, so I wanted to know: shown a string on its own, the way a reviewer sees it, can a model tell a shipped defect from its native fix?
The benchmark has 29 real pairs: the string as it shipped, and the version a native speaker would ship. Brand names are masked. The model sees each half separately, with the product type, where the text appears and the English source when there is one, and answers ok or error.
The scoring is the part I care about. A pair counts only when the model flags the shipped string and leaves the fix alone. A model that flags everything scores zero, and so does one that approves everything. Reviewers that cry wolf are useless in practice, and a plain accuracy number would hide them.
The defects cover ten types: 8 untranslated leftovers, 5 calques, 4 agreement errors, 3 number formats, 2 each of wrong word sense, misspelling and broken grammar, and one each of a broken encoding, an unidiomatic phrase and English-style capitalisation.
Models Tested
I ran the task on the 18 models Kaggle Benchmarks offered me, from Google, Anthropic, OpenAI, xAI, Alibaba, DeepSeek, Zhipu and the open Gemma and gpt-oss families, big and small, because a localisation team picks a reviewer by cost as much as by quality.
Not all of them produced answers. In the first runs, with all models in parallel, seven returned API errors on most calls. I re-ran those seven one at a time on Oct 1. Five then answered cleanly, while three open models still hit rate limits and one was not found, so I report those four as not measured rather than as zero. A run counts only when at most 2 of its 58 answers were unreadable. That leaves 14 models and 21 clean runs.
Findings
Pairs fully right, bugs caught and fixes left alone, out of 29. Where a model has two clean runs, both are shown.
| Model | Pairs fully right | Bugs caught | Fixes left alone |
|---|---|---|---|
| Gemini 3.8 Flash | 26 · 25 | 27 · 26 | 28 · 27 |
| Gemini 3.7 Flash | 26 · 23 | 27 · 27 | 28 · 25 |
| Gemini 3.5 Flash-Lite | 26 · 25 | 26 · 27 | 29 · 27 |
| GPT-6 Astra | 26 | 27 | 28 |
| Gemma 4 31B | 23 · 24 | 25 · 27 | 26 · 26 |
| Gemini 3.1 Pro (preview) | 23 · 24 | 27 · 27 | 25 · 26 |
| Claude Opus 5 | 24 | 26 | 27 |
| Claude Sonnet 5 | 23 | 26 | 26 |
| Claude Haiku 4.5 | 22 | 24 | 27 |
| GPT-5.5 | 22 | 29 | 22 |
| Gemini 2.5 Flash | 21 | 24 | 26 |
| GPT-5.4 mini | 19 · 19 | 24 · 24 | 23 · 23 |
| GLM-5 | 19 | 26 | 22 |
| GPT-5.4 nano | 18 · 16 | 26 · 26 | 20 · 19 |
Catching the bug is the easy half. Every clean model caught 24 or more of the 29 shipped defects. What separated them was the other half: leaving a correct native string alone. GPT-5.4 nano catches as many bugs as the leaders and then flags a third of the native fixes too. GPT-5.5 is the extreme case: the only model to catch all 29 bugs, and one that flagged 7 of the 29 native fixes on the way. A reviewer like that sends a translator to "fix" good Italian, which is worse than no reviewer.
Numbers are the blind spot. Across the 21 clean runs, models caught every untranslated leftover and broken encoding, 103 of 105 calques and 83 of 84 agreement errors. They caught the three number-format errors only 25 times in 63, and only GPT-5.5, the model that flags everything, caught all three. All three are the English "+" glued to a number: 6+ piattaforme, 15.000+ computer, and oltre 100+ siti, which says "more than" twice. Italian writes oltre 6. The strings look like numbers, so models wave them through.
The small model tied the big ones. Gemini 3.5 Flash-Lite, the cheapest Gemini on the list, left all 29 native fixes alone in one run and tied the top score. Gemini 3.1 Pro did not beat the Flash models, and Claude Opus 5 scored two pairs above Haiku 4.5, inside the noise below.
One run is not a result. The same model moved by up to three pairs between runs (Gemini 3.7 Flash scored 26 and then 23; my first notebook run gave 25). On a 29-pair benchmark, a one- or two-pair gap between models is noise; I would not rank models on it.
What it changed about how I think
I used to judge an AI reviewer by what it catches. After this I judge it by what it leaves alone. Recall is cheap: the weakest clean model still caught 24 of 29 bugs. The cost sits in the false alarms, because each one sends a translator off to defend a correct string.
I also stopped trusting "give the model the style guide" as the fix. I tested that on a smaller follow-up for another DEV challenge: the same 29 pairs split so the 12 test pairs never reached the guide, a different model (Gemini 3.1 Flash-Lite), and a real Italian style guide served to it as a knowledge base. On its first run the style guide made it worse, 4 of 12 held-out pairs right against 7 without it, because the model applied rules whose trigger was not in the string. It took two rounds of fixes to bring the grounded version level with the plain one, 9 of 12 each. In that test, the rules raised false alarms long before they lowered anything.
What I'd measure next: more pairs per defect type, especially numbers and idiom, so each type gets its own reliable score, and whether a rule written for exactly the "+" pattern closes the number blind spot without new false alarms.
My Benchmark
The benchmark and its leaderboard are public on Kaggle: https://www.kaggle.com/benchmarks/giuseppecastelluccio/italian-ui-strings-shipped-vs-fixed
The task behind it, with its notebook and every model's run: https://www.kaggle.com/benchmarks/tasks/giuseppecastelluccio/italian-ui-strings-shipped-vs-fixed
How this was made
The strings and their corrections come from my QA audits, made before the challenge; I checked every correction again as a native speaker for this benchmark. The task code was written during the challenge with an AI coding assistant (Claude Code): I set the scoring rule and the checks, it drafted the code, and I reviewed and ran it. The two charts were drawn from the same run results with a short matplotlib script. A stub model that always answers right scores 1.0 and one that flags everything scores 0.0.














