This article is a submission for the Kaggle Benchmarking Challenge.
Every data-science tutorial makes a claim: "I split first so nothing leaks," "cross-validation proves it generalizes," "99% accuracy." The code below the claim does not always agree. So I asked one question: when an AI model reads a real Kaggle notebook, does it trust the caption or read the code?
Blog vs Bytecode is a benchmark built entirely from real, public Kaggle notebooks. I pulled 13 popular kernels, cropped 34 short snippets and paired each with the claim its author (or a plausible write-up) makes about it. Half are honest and sound. Half hide a real methodology flaw. The model returns one verdict, OK or PROBLEM, plus its reason. The first thing the benchmark caught was not a model. It was the harness.
The benchmark
34 items. The design is deliberate:
- Real code, not toy snippets. 28 items are real notebook code (verbatim or trimmed); 9 take real code and inject exactly one controlled flaw. Every item names the kernel it came from.
- Balanced 17 PROBLEM / 17 OK, so a model that reflexively yells "PROBLEM" scores near 50% and gets caught.
- Six flaw families: data leakage, train/test contamination, misused cross-validation, metric mismatch, target-leakage/snooping and claim-vs-code divergence.
-
Matched adversarial pairs that separate understanding from pattern-matching. A RobustScaler inside a pipeline is fine; the same scaler fit once before cross-validation leaks.
KFold(shuffle=True)is fine on independent house-price rows but breaks a time series.np.log1pbefore the split is fine because it fits no parameters. - Genuine flaws from famous notebooks. A top fraud-detection notebook resamples before cross-validating and reports "accuracy" on data that is 0.17% fraud and even labels its own cell "Overfitting Case." A top house-prices notebook fills missing values with a median computed over train and test combined. These are real, not staged.
Three cuts, not one number: overall accuracy, PROBLEM-recall (does it catch the flaw) versus OK-specificity (does it leave sound code alone), plus accuracy on the matched pairs.
Surprise 0: the harness was lying
Kaggle runs the models for you. For several, the run finished but no answer came back: gpt-oss-120b returned nothing on 33 of 34 items, DeepSeek-R1 on 23, Qwen3-Next on 33. Grading those blanks as wrong would rank strong models near zero. A benchmark that does not audit its own capture measures the plumbing, not the model, so the scorer drops uncaptured answers rather than counting them wrong. Everything below is only models whose answers were captured.
Surprise 1: every model catches the flaw, then cries wolf on the clean code
On real notebooks the split is stark. Every captured model scores 100% on PROBLEM-recall. They catch the leak, the resample-before-CV, the accuracy quoted on 0.17% fraud. Where they fall apart is the opposite error: flagging sound code as broken.
Look at the bottom of that chart. Grok with reasoning off cleared zero of the clean notebooks: it called every sound snippet a PROBLEM. Claude Haiku cleared 27%. The expensive mistake on real data is not gullibility. It is the false alarm.
Read the models' own words
The sharpest case is np.log1p applied to the target before the split. It is the textbook exception to "fit transforms inside the fold," because log1p fits no parameters: each row's value depends only on itself. Nothing leaks. Four models flagged it anyway.
Claude Haiku 4.5 argued itself into the false alarm:
Although np.log1p is indeed a fixed elementwise transform with no parameters ... the broader methodological issue is that the transformation is applied before any cross-validation split.
It states the correct fact, then overrides it. Claude Opus 5, on the same snippet, held the line:
log1p is a deterministic, parameter-free transform: each row's transformed value depends only on that row, so no statistic is shared across the fold boundary. This is the standard exception to the 'fit transforms inside the fold' rule.
That gap, knowing the rule versus knowing its exception, is most of the leaderboard.
When the flaw is real, the models are sharp. On the house-prices notebook that fills LotFrontage with a neighborhood median computed over train and test combined, every model caught it. Gemini 3.8 Flash:
The code computes the median LotFrontage per neighborhood on 'all_data' (which concatenates both train and test rows), thus leaking test set information into the imputation rather than computing medians solely from the training rows as claimed.
Surprise 2: reasoning is worth 40 points, all of it specificity
Grok 4.20 scores 94% with reasoning on and 52% with it off. The 42-point gap is entirely on the clean code: OK-specificity falls from 88% to 0%. With reasoning off the model stops reading what a line computes and pattern-matches "this is a notebook, notebooks have bugs, PROBLEM." It flagged the parameter-free log1p too.
Surprise 3: the small model trusts the prose
Gemma 4 31B is the only model that misses the flaws: 35% recall against everyone else's 100%. It reads the confident claim, agrees, then walks past the leak. So the bias flips with scale. Big models over-flag sound code while the small model under-flags real bugs. Same benchmark, opposite failure.
The leaderboard
Scored on the items the proxy captured; models it dropped on most items are left off.
| Model | Accuracy | Catches flaws | Clears clean code |
|---|---|---|---|
| Claude Opus 5 | 94% | 100% | 88% |
| Grok 4.20 (reasoning) | 94% | 100% | 88% |
| GLM-5 | 93% | 100% | 85% |
| Claude Sonnet 5 | 93% | 100% | 83% |
| Gemini 3.1 Pro | 93% | 100% | 82% |
| Gemini 3.7 Flash | 91% | 100% | 82% |
| Gemini 3.8 Flash | 88% | 100% | 76% |
| Gemini 3.5 Flash-Lite | 88% | 100% | 75% |
| Claude Haiku 4.5 | 62% | 100% | 27% |
| Grok 4.20 (no reasoning) | 52% | 100% | 0% |
| Gemma 4 31B | 47% | 35% | 59% |
Honest limits
- Two backends. Kaggle's Model Proxy captured most models; a few (DeepSeek-R1, gpt-oss, some Qwen) it dropped, so those are omitted rather than scored on a handful of items. The numbers above are on captured answers only, so a couple of rows are over a subset.
- Injected versus real. 9 of the 17 PROBLEM items are one controlled flaw injected into real code; the rest, plus every OK item, are real notebook code unchanged. The data marks which is which.
- Verdict grading. Grades are on OK versus PROBLEM. The written reason is captured too. On the top models it names the right flaw almost every time.
What I would measure next
- Localization: not just OK or PROBLEM but which line, scored against the known offending line.
- A confidence axis, to see whether the over-flaggers are also the least calibrated.
- More time-series and NLP notebooks, where the leakage is subtler than a misplaced scaler.
Reproduce it
The benchmark is public on Kaggle: kaggle.com/benchmarks/tasks/zkasuran/blog-vs-bytecode. Every item, verdict and grader is open. Each model's answer is captured verbatim, so you can read exactly where it went wrong.
Sources and disclosure
Every snippet is derived from a real public Kaggle notebook. Sources include the credit-card-fraud notebook by janiobachmann, the stacked-regressions and house-prices notebooks by serigne and pmarcelino, the Titanic notebooks by startupsci and arthurtok, the IMDB and disaster-tweet NLP notebooks and several Kaggle Learn exercises. Injected flaws are marked in the data and each item names its source kernel.
AI (Claude) helped build the benchmark and draft this post. The design, the items, the grading and the analysis were verified by the author, including reading every snippet and every raw model output.














