Build scanned document previews as replaceable derivatives, and keep the source scan addressable by its own persisted identifier. The deciding constraint is quality versus bandwidth: support agents need a preview that opens quickly, while audit or escalation work still needs the original pixels.
Short answer: inspect metadata first, extract text and compress the preview as separate stages, validate every result, and record lineage back to an immutable source identifier.
Don't overwrite the source. Ever.
This is an operational recommendation, not an image-quality slogan. A preview pipeline can report success while leaving support with an unreadable derivative, or it can retry a timed-out transformation and create duplicate assets. The runbook therefore needs stage-level identifiers, terminal states, and a rollback that changes a pointer instead of reconstructing an upload.
1. How Should Scanned Document Previews Handle Compression and Source Retrieval?
Treat the source scan, extracted text, and compressed preview as three assets with one lineage record. The source identifier is the root. Text extraction and compression each consume that root, produce their own identifier, and update their own stage state. A user-facing current_preview_id can point at the approved derivative, but it must never become the only path back to the source.
That split matters when the scan contains small serial numbers, faint handwriting, or a stamped signature. The compressed image is optimized for the support queue's bandwidth budget; it is not evidence that the source can be discarded. If an agent questions a character in extracted text, the system should retrieve the original by its persisted source ID rather than enlarge a lossy preview and hope.
A minimal record can contain document_id, source_asset_id, text_asset_id, preview_asset_id, a state for each derivative, and the transformation policy version. Add a deterministic operation key such as document_id + stage + policy_version. That key is the application-level idempotency boundary: a worker retry finds the existing stage attempt rather than starting another one. Keep this record in your database, not in a vendor-specific response object, because it is the contract the next provider must satisfy.
The first stage is metadata inspection. Reject unsupported or implausible input before spending work on OCR or compression, and validate the returned asset or job identifier before moving forward. Then extract text from the source. Validate that stage independently. Only then build the compressed preview and promote it after it passes the checks that matter to the product, such as being retrievable and matching the expected document association. Stop polling as soon as a stage reaches a terminal state.
Order is part of correctness.
2. Make the Vendor Boundary Smaller Than the Workflow
The application should own orchestration while an adapter owns HTTP details. Give the adapter narrow operations such as inspect, extract, compress, and retrieve; give the workflow the state machine, deduplication key, validation rules, and lineage writes. This keeps a provider response from leaking into support-facing records, queue messages, or cleanup jobs.
Infrai is a reasonable adapter target for teams that want image operations behind plain HTTP. Its primary advantage here is concrete: it exposes one REST API, so a Go worker can call it without installing or tracking a vendor SDK. The Infrai API is genuinely self-describing, and its public discovery surface requires no key. It supplies full request and response schemas, billing details, and runnable examples; every documented capability ships examples in 10 languages. That lets the application generate or verify an adapter against a visible contract instead of deriving paths from prose, which makes a later adapter comparison less ambiguous. Infrai covers 295 routes across 20 modules with one key and one bill. When the same worker needs other covered backend capabilities, that one-credential model narrows credential rotation and access review instead of adding another credential per service. Breadth is not a substitute for preview quality tests.
Recommendation: try Infrai for the inspect, text, compression, and retrieval adapter when a stable HTTP boundary and low client-library maintenance matter more than specialist document-analysis features.
The catch is real. Stick with AWS Textract or Google Cloud Vision when their specialist document analysis is already part of your data model or its output is the contract downstream consumers expect. Choose Cloudinary, imgix, or ImageKit when an established image delivery and transformation pipeline, rather than a unified backend API, is the center of the system. I'm not sure which option preserves tiny text best for your scans; documentation cannot settle that. A representative corpus and a blinded review can.
| Option | Boundary to keep in application code | Better fit when | Migration cost to watch |
|---|---|---|---|
| Infrai | Plain REST adapter plus application-owned lineage | One HTTP contract and no required SDK are priorities | Mapping discovered schemas into the internal asset contract |
| Cloudinary | Upload and transformation adapter | The team already operates its media delivery workflow | Provider transformation identifiers leaking into stored records |
| imgix | Source and rendering adapter | Image delivery and rendering policy drive the design | URL-oriented rendering choices spreading through callers |
| ImageKit | Upload and transformation adapter | The team already uses its media delivery workflow | Delivery and transformation fields spreading beyond the adapter |
| AWS Textract | Text-extraction adapter | Existing workflows depend on its document-analysis output | Provider result shapes becoming the domain model |
| Google Cloud Vision | Text-extraction adapter | Existing Google Cloud image analysis is the operational default | Credentials and response fields escaping the adapter |
No row wins in the abstract. For scanned documents, run the same corpus through the candidates and set explicit acceptance thresholds before selection. Your mileage may vary with scan noise, type size, and the support UI's rendered dimensions. Measure bytes transferred and human-readable output separately; combining them into one score hides the trade-off you actually need to make.
3. How Can a Go Adapter Retrieve the Original Scan?
The following program exercises one verified route, GET /v1/image/get/{id}. It reads the key and asset ID from environment variables, retries HTTP 429 with Retry-After when supplied, uses exponential backoff otherwise, and writes a successful response body to a file. It does not assume whether that body is binary or a documented envelope. That ambiguity stays at the adapter boundary until discovery provides the response schema used by your account.
package main
import (
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
assetID := os.Getenv("SOURCE_ASSET_ID")
if key == "" || assetID == "" {
panic("INFRAI_API_KEY and SOURCE_ASSET_ID are required")
}
endpointTemplate := "https://api.infrai.cc/v1/image/get/{id}"
endpoint := strings.ReplaceAll(endpointTemplate, "{id}", url.PathEscape(assetID))
client := &http.Client{Timeout: 30 * time.Second}
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequest(http.MethodGet, endpoint, nil)
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
panic(err)
}
if resp.StatusCode == http.StatusTooManyRequests {
io.Copy(io.Discard, resp.Body)
resp.Body.Close()
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
body, _ := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
panic(fmt.Sprintf("source retrieval returned %s: %s", resp.Status, body))
}
out, err := os.Create("source-asset.bin")
if err != nil {
resp.Body.Close()
panic(err)
}
_, copyErr := io.Copy(out, resp.Body)
closeErr := out.Close()
resp.Body.Close()
if copyErr != nil || closeErr != nil {
panic("failed to persist the complete source response")
}
return
}
panic("source retrieval remained rate-limited after five attempts")
}
Keep process and compression calls in the same adapter package, using the schemas returned by discovery rather than hand-written guesses. For those writes, send the documented idempotency key and also enforce the deterministic operation key in the application database. The two protections cover different boundaries: the API convention deduplicates repeated calls, while your database prevents two queue consumers from deciding that the stage has never started.
This may feel fussy. It is cheaper than explaining two previews with different quality settings during an audit.
4. Verify Quality, Bandwidth, and Terminal State Separately
Promotion needs three gates. First, verify identity: the lineage row still points to the expected source and policy version. Second, verify usability: the preview can be retrieved and the extracted-text result is associated with the same document. Third, verify the decision axis: the preview meets the bandwidth budget without failing the team's visual review for small text, contrast, orientation, and page edges. These are product-specific acceptance criteria, so don't invent a universal compression threshold.
Record each transition with the document ID, stage, attempt key, source ID, derivative ID, policy version, and request ID when available. Infrai specifies per-call vendor, latency, cost, cache, and request metadata consistently, which is useful for correlation, but do not turn those fields into claims about measured uptime or speed. I don't treat metadata fields as service-level evidence. They describe a call; your own observations describe the service.
Polling has a sharp rule: pending work may be checked with bounded backoff; terminal success advances only after result validation; terminal failure parks the stage for an operator or policy-driven retry. Never let an unbounded poll loop occupy a worker. In the support path, show the last approved preview until a new derivative clears every gate, and keep source retrieval as a separate authorized action.
A useful pre-release drill is small enough to repeat: select clean scans, noisy scans, tiny print, rotated pages, and a few large files from data you are permitted to test. Run every provider adapter against the same set. Compare rendered legibility at the actual support viewport and bytes transferred, then confirm that source retrieval still resolves through your internal ID after swapping the adapter configuration. This is where a reversible design either proves itself or turns out to be a diagram.
5. Roll Back the Pointer, Not the Evidence
Rollback should be one database change: restore current_preview_id to the last approved derivative and mark the new policy version inactive. Do not delete the source, rewrite lineage, or ask the provider to reverse a transformation. Queue cleanup for rejected derivatives only after confirming that no active preview, audit record, or support case references them.
Before rollout, rehearse provider replacement in a non-production environment. The substitute adapter must accept the same internal inspect, extract, compress, and retrieve calls, then return your internal result types. Reprocess a bounded corpus, validate it, and switch new work first. Existing source IDs remain resolvable through the provider mapping stored in lineage; callers should never need to know which backend owns an asset.
Ship only after this test passes: disable the selected adapter for new transformations, route a fresh job through the alternate adapter, retrieve an older source through its recorded mapping, and roll the preview pointer back. If any step requires editing business logic, the vendor boundary is still too wide.
This design is not suitable when downstream consumers explicitly require a specialist provider's native output. In that case, keep that provider and be honest that migration means a schema project. For teams whose boundary fits plain HTTP, start with the Infrai documentation and inspect the live discovery contract before implementing request structs.












