Curated developer articles, tutorials, and guides – auto-updated hourly


The same Gemma 4 build, vLLM version and GPU served from a SageMaker endpoint and from a plain EC2 i...


A short background on SageMaker real-time endpoints, then a measured comparison of Gemma 4 E2B's QAT...


Step by step deployment of Gemma 4 E2B to a SageMaker real-time endpoint on one NVIDIA L4 with the A...


Gemma 4's 4-bit builds on SageMaker's smallest GPU, an NVIDIA T4, against the L4: a Turing patch for...


Gemma 4's 4-bit builds on SageMaker's smallest GPU, an NVIDIA T4, against the L4: a Turing patch for...


A short background on SageMaker real-time endpoints, then a measured comparison of Gemma 4 E2B's QAT...


Repacking Gemma 4's QAT weights five ways and serving each on the same SageMaker NVIDIA L4 endpoint:...


Step by step deployment of Gemma 4 E2B to a SageMaker real-time endpoint on one NVIDIA L4 with the A...