Curated developer articles, tutorials, and guides – auto-updated hourly


A nine-tool survey of local-model interfaces on one GPU: what worked, what silently...


Pros and cons of both tools provided by Bob! Introduction I’ve been reading and seeing...

A llama.cpp release note does not prove that Ollama can use the feature. I trace the active runtime ...


LLM inference in macOS VMs collapses to 12.63 tok/s because the guest reports GPU family 5 and llama...


Originally published on andrew.ooo — visit the original for any updates, code snippets that aged...


The most widely used local AI inference engine published its first semantic version tag on August 17...