When a media upload creates a responsive thumbnail, storage and cache cost depend on one small architectural distinction: bytes and work state are different resources. Short answer: retrieve an asset only after its representation is ready; poll an asynchronous job status when the representation is still being produced. Returning one shape for both paths makes clients cache the wrong thing and turns a harmless refresh into more work.
I keep two identifiers in the upload record: an immutable asset key and a job identifier. The key answers “where are the bytes?” The job answers “is the transformation finished?” That split stays useful in a Python worker, a Node.js edge handler, or a notebook experiment that later becomes production code.
How should asset retrieval differ from asynchronous work state?
An asset retrieval response is a representation request. It should carry a stable cache identity, a media type, and bytes (or a redirect to bytes). A job-status response is a control-plane observation. It should carry state such as queued, running, succeeded, or failed, plus an optional output key. The status document is small and short-lived; the thumbnail is immutable and cacheable for much longer.
That distinction prevents a common bug in client code: treating 202 Accepted as if it were a thumbnail. A 202 says the request was accepted for processing; it does not provide a representation to decode. The client can poll the status endpoint with backoff, then switch to the retrieval endpoint once an output key exists. A 404 for the output during running is not a useful contract; the status response should explain readiness directly.
Here is the small state machine I use in tests. It makes the cache boundary explicit without tying the design to a vendor API.
from dataclasses import dataclass
from typing import Literal
State = Literal["queued", "running", "succeeded", "failed"]
@dataclass(frozen=True)
class ThumbnailJob:
asset_key: str
job_id: str
state: State
output_key: str | None = None
def retrieval_allowed(self) -> bool:
return self.state == "succeeded" and self.output_key is not None
The test cases are intentionally boring: every non-success state must remain a status response, and only succeeded can mint a retrieval URL. Boring contracts are cheap to evaluate and easy to observe.
Where do cache keys and storage costs diverge?
Responsive media multiplies objects quickly. One original upload might produce widths 320, 640, and 1280, each in WebP and JPEG. If a status document is cached for five minutes while a thumbnail is cached for a day, the traffic pattern is predictable: many tiny control-plane reads, fewer large data-plane reads. Combining them forfeits that leverage because a changing status invalidates the object response. The difference shows up in a cost report as well as in latency: a status poll is measured in a few headers and a short JSON body, while a miss for a 1280-pixel image can pull hundreds of kilobytes from origin. A cache key that includes job state makes every transition a miss; a key that names the immutable output lets the same bytes serve mobile, desktop, and later replays until the retention policy says otherwise. That is why I put width, format, and encoder version in the object identity, then keep viewer-facing URLs stable through a small manifest or redirect layer. It takes more bookkeeping up front, but it makes the expensive part observable instead of accidental.
Use content-addressed or versioned output keys, such as photo/8f2c/w640-v3.webp, rather than overwriting photo/8f2/w640.webp. Immutable keys let a CDN retain a hit while a later upload or encoder change creates a new version. Keep the job record mutable, but keep the output object immutable. That is the practical line between status and retrieval.
The storage bill is not the only signal. Measure cache-hit ratio by width, bytes served from origin, status polls per completed job, and the median time from upload to first usable thumbnail. I also record duplicate-job rate: retries that enqueue the same asset are a direct cost leak. Your mileage may vary when originals are very large or viewers request unusual widths, so measure a representative week before choosing retention windows.
Keep it boring.
What failure modes appear when the two paths share a contract?
The first failure is stale success. A client caches a response that says running, then never asks again because the cache metadata looked like an image response. The second is a thundering herd: every browser retries the upload route instead of polling a cheap status document. The third is accidental disclosure, where a status endpoint returns a long-lived signed URL before access checks run on the asset path.
I prefer explicit transitions and idempotency. The upload request can accept an idempotency key derived from the original asset digest and requested widths. A worker may run twice, but it should publish one versioned output and one terminal status. If a queue is delayed, the status remains queued; it does not masquerade as an empty image.
Retry policy belongs to the status client. Exponential backoff with jitter protects the queue, while a Retry-After value gives the server a voice. Retrieval failures follow media semantics instead: validate the media type, honor range requests where supported, and return a cacheable response only when the object is complete.
A decision rule for notebook-to-production systems
Start with one fixture: an upload, three target widths, and a deliberately slow worker. In the eval harness, assert that each poll is smaller than each image response, that a completed job never changes its output key, and that a second upload with the same digest does not create duplicate objects. Then replay the fixture with a cold cache and a warm cache; compare origin bytes, not just request counts.
This approach is not suitable when the transformation must be streamed live, such as an interactive video preview. In that case, a session or stream protocol is a better fit than a finite job record. Stick with a single synchronous retrieval call when images are tiny and processing is guaranteed inside the request budget; adding a queue would add state without buying useful cache behavior.
The trade-off is operational surface area: two endpoints, a worker, expiry rules, and metrics. I accept that cost for media systems because it keeps expensive bytes stable while cheap status changes rapidly. The important part is the contract, not the brand of object store or queue behind it.













