Curated developer articles, tutorials, and guides – auto-updated hourly


A preliminary report concluded that models under 3B parameters cannot chain exploits. A later tourna...


I built a benchmark to rank language models on spotting liars, then ran 522 games of Mafia through i...


Benchmark scores are the industry's favourite number and one of its least honest. Here is why the le...


Dan Luu just published "The Benchmarkpocalypse," a devastating analysis of why AI benchmarks are...

Frontier models are dropping factual accuracy to gain reasoning efficiency. The tradeoff is starting...


Five side channels that let a coding agent re-derive the fact your memory eval is meant to measure: ...


xAI's new flagship matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index — 500K context,...


MiniMax H3 and the Free-Tier Trap: Run a 12-Prompt Gate Before You Switch 12 prompts....


V4 Pro 0813 scores 53 on Artificial Analysis at $0.06 per task, 5x under GLM-5.2. MIT weights are no...


GLM 5.3 runs on the Coding Plan with 1M context, 128K output, Terminal Bench 3.0 improved to 28.3, w...


Cold starts and memory decide your infra bill. WebSockets decide your tail latency. Here’s how Bun a...


A systematic evaluation of seven frontier models on 36 long-horizon research tasks found that agents...


VibeWorlding tests whether multimodal agents can turn a plain request into an interactive 3D scene e...


Terminal-Bench 3.0 launched with 74 tasks across seven domains, and the top agent scores 34.4 percen...


A new routing framework that chooses a different model for each request outperformed the strongest s...


A new benchmark replaced scripted evaluation with an AI agent pursuing long-horizon goals inside gen...


Alibaba's Qwen3.8-27B shipped with the same coarse architecture as Qwen3.6-27B, prompting accusation...


StateM reaches 95.3 percent raw accuracy on Terminal-Bench 2.1 across 445 trials without changing an...


A benchmark of 17,886 crisis videos found that a single round of ordinary resharing dropped the best...