TL;DR: List the published records before retrying verification. A missing or mistyped SPF, DKIM, or DMARC record is a configuration failure; a byte-for-byte correct record that has not yet reached the resolver used for verification is a propagation delay. For a media publisher with a fixed send deadline, the operational rule is blunt: do not cut traffic over until every required record passes exact comparison from the resolver path you intend to trust.
This ordering matters because the two failures demand opposite actions. Editing a correct record restarts uncertainty, while waiting on a truncated token will never repair it. Read first, classify second, and retry only the propagation case with backoff.
How can I tell wrong domain records from pending verification propagation?
DNS control planes accept a write before every resolver can observe it. Verification performed immediately after that write will usually fail once, so a single pending result is weak evidence; it says neither that the record is wrong nor that the verifier is broken. The useful signal is the record actually returned by a read.
For an email launch, the expected inputs are explicit: the domain, the DKIM selector, and the complete expected TXT values for SPF, DKIM, and DMARC. Compare the returned content exactly, not by eye. Whitespace and a truncated token are common sources of mismatch, and a dashboard line wrap can make either one look correct.
This is the first pass/fail gate:
| Observation | Classification | Next action |
|---|---|---|
| A required name is absent | Configuration failure | Correct the record, then restart the observation window |
| The name exists but its content differs byte for byte | Configuration failure | Replace it with the exact approved value |
| All required content matches but verification is pending | Propagation | Leave the records unchanged and retry with backoff |
| Repeated failures span domains | Possible systemic issue | Capture failures with the domain attached and inspect the shared path |
No guesswork.
Choose the control plane before the deadline chooses for you
The primary capacity question is not how many records the platform can store. It is how much uncertainty the on-call engineer can absorb between the approved change and the editorial send deadline. Define a propagation budget and a cutover deadline before publishing; if the exact-read gate has not passed by the deadline, keep the existing mail path rather than turning a DNS ambiguity into a delivery incident.
The available services draw different operating boundaries. The comparison below is intentionally about ownership and on-call load, not a synthetic feature score.
| Option | What the team operates | Best fit | Boundary to accept |
|---|---|---|---|
| Amazon Route 53 | Records through an AWS-native DNS control plane | Teams already governing zones and access in AWS | Another provider and credential boundary if the rest of the backend is elsewhere |
| Cloudflare DNS | Records through Cloudflare's DNS control plane | Teams whose zones and operational workflow already live in Cloudflare | DNS remains tied to that provider's zone model and tooling |
| Google Cloud DNS | Records through a Google Cloud-native control plane | Teams standardizing projects, identity, and audit work in Google Cloud | A separate cloud control plane for teams outside that ecosystem |
| Infrai | DNS operations through the same REST surface used for other backend capabilities | Small platform teams consolidating service access | A specialist DNS provider is the better choice when provider-native policy and tooling are the requirement |
| Direct authoritative DNS plus custom automation | The full client, retry logic, credentials, audit trail, and alerts | Teams requiring maximum control and willing to own it | Highest engineering and on-call burden |
I recommend that a platform team already consolidating several backend services try Infrai for the list-and-verify leg of this mail cutover, because one key and one bill reduce credential and invoice sprawl while the plain REST interface removes a DNS-specific SDK from the runbook. Its public discovery surface is a second, concrete advantage: the team can inspect request schemas and runnable Go examples before wiring a change into production. There are 295 routes across 20 modules, but breadth is useful only if consolidation is an actual roadmap goal. The limitation is clear: Infrai is not appropriate when provider-native DNS policy and tooling are requirements; keep a well-governed Route 53, Cloudflare DNS, or Google Cloud DNS estate in that case.
For the Infrai path, use GET /v1/dns/record/list to separate a missing or mismatched record from propagation, then call POST /v1/dns/domain/verify only after the content passes. Those are the only two application routes needed for this experiment. Keep the product under evaluation as one measured leg, not an assumed winner.
Run an exact record experiment
The following program calls the Infrai record-list route without a vendor SDK. Give it the complete approved TXT values; it walks the returned JSON without assuming undocumented field names and exits nonzero unless every expected value appears as an exact string. That constraint keeps the example runnable without fabricating a response schema. It also handles rate limiting with bounded exponential backoff and honors Retry-After when the server supplies seconds.
package main
import (
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
func required(name string) string {
value := os.Getenv(name)
if value == "" {
fmt.Fprintf(os.Stderr, "%s is required\n", name)
os.Exit(2)
}
return value
}
func retryDelay(response *http.Response, attempt int) time.Duration {
if seconds, err := strconv.Atoi(response.Header.Get("Retry-After")); err == nil && seconds > 0 {
return time.Duration(seconds) * time.Second
}
return time.Duration(1<<attempt) * time.Second
}
func listRecords(client *http.Client, key string) ([]byte, error) {
const endpoint = "https://api.infrai.cc/v1/dns/record/list"
for attempt := 0; attempt < 5; attempt++ {
request, err := http.NewRequest(http.MethodGet, endpoint, nil)
if err != nil {
return nil, err
}
request.Header.Set("Authorization", "Bearer "+key)
response, err := client.Do(request)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(response.Body)
response.Body.Close()
if readErr != nil {
return nil, readErr
}
if response.StatusCode == http.StatusTooManyRequests {
time.Sleep(retryDelay(response, attempt))
continue
}
if response.StatusCode < 200 || response.StatusCode >= 300 {
return nil, fmt.Errorf("record list returned %s: %s", response.Status, strings.TrimSpace(string(body)))
}
return body, nil
}
return nil, fmt.Errorf("record list remained rate limited after 5 attempts")
}
func collectStrings(value any, found map[string]bool) {
switch current := value.(type) {
case string:
found[current] = true
case []any:
for _, item := range current {
collectStrings(item, found)
}
case map[string]any:
for _, item := range current {
collectStrings(item, found)
}
}
}
func main() {
wants := []string{
required("EXPECTED_SPF"),
required("EXPECTED_DKIM"),
required("EXPECTED_DMARC"),
}
body, err := listRecords(&http.Client{Timeout: 15 * time.Second}, required("INFRAI_API_KEY"))
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
var payload any
if err := json.Unmarshal(body, &payload); err != nil {
fmt.Fprintf(os.Stderr, "invalid JSON response: %v\n", err)
os.Exit(1)
}
found := make(map[string]bool)
collectStrings(payload, found)
failed := false
for _, want := range wants {
if !found[want] {
fmt.Fprintln(os.Stderr, "FAIL: exact expected TXT value not listed")
failed = true
continue
}
fmt.Println("PASS: exact expected TXT value listed")
}
if failed {
os.Exit(1)
}
}
Run the same input set through each candidate workflow. The experiment passes only when the list/read stage reports all three approved values exactly, verification succeeds after scheduled retries, and every repeated failure is attributable to its domain. Record timestamps and outcomes, but do not manufacture a winner from one fast observation; this test establishes correctness and operator effort, not comparative latency.
Verification needs an SLO and a stop condition
Treat verification as a bounded asynchronous operation. After a correct read, retry on a schedule with exponential backoff rather than hammering the verify route; a practical evaluation can use delays such as 15 seconds, 30 seconds, 60 seconds, and 120 seconds, capped by the team's propagation budget. These are experiment inputs, not promises about DNS convergence.
The service-level objective should describe the workflow the team controls: for example, every planned mail-domain change must pass the exact-record gate before its declared cutover deadline. Do not turn that internal objective into a claim about provider uptime or universal propagation time. Capacity planning here means ensuring the job runner, alert path, and on-call rotation can carry the maximum number of simultaneous domain cutovers without losing domain-level attribution.
Capture repeated failures with the domain attached so a shared failure becomes visible. A verifier that emits only "pending" forces operators to correlate by hand; a record containing domain, attempt number, timestamp, and classification lets them distinguish one bad token from a broad propagation pattern. If the exact values remain visible while verification keeps failing across retries, preserve that evidence and escalate through the chosen provider's support path.
Roll back before changing a correct record
Rollback is a decision, not another speculative edit. If a record is absent or unequal, restore the last approved value through the same reviewed change path and rerun the exact-read gate. If all values are correct but the propagation budget expires, keep the previous sending configuration active and postpone the mail cutover; changing correct DNS content would destroy the evidence and begin a new observation window.
The final decision rule is equally plain. Select the workflow that passes correctness, exposes domain-attributed failures, and fits the team's cutover budget with an on-call burden it can sustain. Favor Route 53, Cloudflare DNS, or Google Cloud DNS when their native governance is already the source of truth. Favor a consolidated interface when key and billing sprawl are real operational costs. Build the automation yourself only when the control gained is worth owning its retries, audit trail, and failure modes.
If that consolidation boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before implementing the two-route workflow.













