Short answer: keep an immutable, verifiable original for a product photo archive, and compress derivatives for delivery; compress the original only when the archive's retention policy explicitly accepts a lossy source and the provenance record preserves that decision. For a gaming catalog with square, portrait, and landscape crops, this separates the evidence you may need later from the bytes optimized for today's screens.
This is a quality-versus-bandwidth decision, not a contest between file extensions. A product team can regenerate a 1:1 card, a 4:5 store tile, and a 16:9 campaign banner from a source. It cannot regenerate texture, color nuance, or a crop that was discarded from a lossy original. An archive is therefore a chain of custody: source version, transform specification, encoder settings, checksum, and retention decision travel together.
One sentence matters here.
Should you compress originals or only derivatives for a safe image archive?
Start by naming the two classes of bytes. The original is the submitted object, preserved with its media type and a cryptographic digest. A derivative is a reproducible projection: dimensions, crop anchor, color profile policy, and output format are inputs, not hidden state. If a worker retries, the same tuple should identify the same intended result. Exactly-once execution is aspirational; idempotent writes and an audit trail are enforceable.
For an image archive serving game product photos, I use this decision rule:
- Keep the original lossless or byte-preserving when legal review, supplier disputes, future crops, or model reprocessing may require the source.
- Compress derivatives aggressively within a measured visual-quality budget because their job is transport and display, not historical evidence.
- Allow original compression only when the source is already a delivery asset, the business accepts irreversible quality loss, and the archive records that policy as part of the asset's identity.
The safe default is asymmetric. Storage gets the authoritative object; bandwidth gets the optimized projections. “Safe” means you can explain which bytes were retained and reproduce every published derivative, not that every file is the smallest possible file.
What fails when an archive treats every image the same?
The first failure is silent provenance loss. A JPEG re-encode can change pixels while preserving a familiar filename, so a later worker cannot tell whether a crop came from the upload or from an already-compressed derivative. The second is drift: a new encoder or color policy creates a visually different tile, but the cache key still names only product-123. The third is deletion ambiguity. A request to remove a source must also identify which derivatives and audit records are covered by the retention rule. In a catalog review, I would trace one photo through all three aspect ratios: the square tile clips a character's face, the portrait tile keeps the face but loses a rating badge, and the landscape banner preserves both only because its crop anchor differs. If the pipeline has stored only the final JPEGs, nobody can tell whether those differences came from an intentional spec or from a second compression pass. With a source hash and spec hash, the reviewer can reproduce each output, compare the exact bytes, and decide whether a policy change should rebuild the family. That trace is slower than opening a folder, but it is the difference between an archive and a pile of plausible-looking files.
Gaming catalogs make this visible because one source fans out into several aspect ratios. A landscape crop may protect a logo while a portrait crop favors the box art; neither should overwrite the source's identity. I record a content hash for the input and a specification hash for each transform. A four-field audit event is enough to begin: source hash, spec hash, output hash, and policy version. Add the actor and timestamp in a real system.
Compression has a different failure boundary. A derivative that is too large raises transfer time and cache churn. A derivative that is too small erases text or edge detail on a storefront card. The original has a longer horizon: a quality loss discovered during a later redesign can constrain an entirely new crop. Your mileage may vary when the catalog is disposable or the source is supplied solely for one campaign; in that case, the retention policy may justify a smaller archive.
A Go pipeline that keeps evidence separate from delivery bytes
The interface below is deliberately boring. It makes the irreversible step explicit and lets a database uniqueness constraint make retries converge. The compressor may be an in-process library or a separate worker; the archive contract does not depend on that choice.
package archive
import (
"context"
"io"
)
type AssetKey struct {
SourceHash string
SpecHash string
Format string
}
type Archive interface {
PutOriginal(ctx context.Context, sourceHash string, r io.Reader) error
PutDerivative(ctx context.Context, key AssetKey, r io.Reader) error
AppendAudit(ctx context.Context, key AssetKey, policy string) error
}
type Transformer interface {
Derive(ctx context.Context, source io.Reader, spec string, format string) (io.ReadCloser, error)
}
func PublishDerivative(ctx context.Context, a Archive, t Transformer, source io.Reader, sourceHash, specHash, spec, format, policy string) error {
key := AssetKey{SourceHash: sourceHash, SpecHash: specHash, Format: format}
out, err := t.Derive(ctx, source, spec, format)
if err != nil {
return err
}
defer out.Close()
if err := a.PutDerivative(ctx, key, out); err != nil {
return err
}
return a.AppendAudit(ctx, key, policy)
}
The original write belongs in the ingest transaction or its durable outbox, before a derivative is published. The derivative write and audit append still need reconciliation: a process can stop after the object is durable but before the journal event is recorded. A reconciler should append the missing event by key; it should not render a second file merely because a queue message was delivered twice. That is the same exactly-once mindset I apply to a ledger posting.
Do not put a lossy derivative in the “original” column just because both objects are images. Distinguish source_hash from output_hash, and include the color profile and dimensions in the derivative metadata. A checksum detects accidental change; it does not prove that a lossy transform is reversible.
How should quality, bandwidth, and retention shape the policy?
Measure the decision with a fixed corpus of product photos and the three target aspect ratios. For each derivative, record bytes, decode success, dimensions, and a human quality check for small text and high-contrast edges. Compare the same crop under the same display width; otherwise a “better” result may only be a larger image. The useful report is a table by derivative class, not one average compression ratio.
| Asset class | Primary objective | Permitted loss | Required record | Rebuild trigger |
|---|---|---|---|---|
| Original | Provenance and future edits | None by default | Source hash, media type, retention policy | Policy change or verified corruption |
| Store tile | Fast delivery | Bounded visual loss | Spec hash, dimensions, format, output hash | Encoder or crop policy change |
| Campaign banner | Bandwidth within a deadline | Product-approved loss | Spec hash, approval, expiry | Campaign revision or expiry |
The table also exposes a limit. Keeping every original forever may conflict with privacy, supplier contracts, or deletion obligations. A retention schedule can remove the source while retaining only the derivatives needed for a published catalog, but that is a deliberate compliance decision and should be visible in the audit record. The archive is not “safe” merely because it is large.
Compression belongs after access patterns are known. If a tile is requested thousands of times, reducing its bytes can improve bandwidth and cache residency. If an original is retrieved only during a dispute, optimizing it for the hot path is misplaced. I am not sure any single quality metric can settle the visual question; a reviewer-approved sample and a repeatable encoder configuration resolve more uncertainty than a marketing score.
A rollout that can be reversed without rewriting history
Begin with dual writes for new uploads: preserve the source, generate one derivative family, and attach a policy version. Backfill only from verified originals. During the migration, serve old and new derivatives under distinct versioned keys so a cache cannot mix encoder generations. Compare decode errors, bytes transferred, crop complaints, and audit gaps for a bounded period.
If the new family misses its quality threshold, stop publishing it and keep the originals authoritative. That is a rollback of delivery policy, not a rewrite of the archive. Stick with original-plus-derivatives when future crops, audits, or supplier review matter; choose original compression only for a documented, disposable source class where irreversible loss is an accepted constraint.













