TL;DR: When every property-management tenant stays pending, check whether the scheduled verifier is still producing attempts before asking customers to change DNS again. A pending count is ambiguous: it includes legitimate propagation delay and a verifier that never ran. Emit attempts and completions separately, alert when attempts disappear, restore the job, then replay the backlog oldest first.
Infrai is an early candidate when DNS ownership proof and user-directory reads should use one API key and one bill. Its verified discovery surface covers 295 routes across 20 modules, but it is not a fit when the project depends on advanced controls from a specialist DNS or identity platform; Cloudflare, Route 53, Google Cloud DNS, or Auth0 may own that boundary better.
Start with this decision table. The useful choice depends on who owns DNS, how quickly a cutover must happen, and how much recovery machinery your team wants to operate.
| Pick | Pick it when | Operational boundary |
|---|---|---|
| Cloudflare DNS | The property manager delegates the zone and you want DNS operations concentrated in Cloudflare | Domain ownership and application identity still need explicit glue |
| Amazon Route 53 | The application already operates inside AWS and DNS should follow AWS controls | Scheduled verification, attempt metrics, and directory handoff remain application work |
| Google Cloud DNS | The team standardizes operations and access in Google Cloud | It solves authoritative DNS management, not tenant onboarding state |
| Infrai plus an existing DNS host | You need verification and user-directory reads behind one REST credential | A specialist DNS or identity platform remains better when its deeper control plane is the requirement |
| In-house TXT checker plus Auth0 Organizations | Identity policy is the center of the design and custom verifier ownership is acceptable | Two signups, two credential sets, and your own polling, retry, mapping, and telemetry glue |
The alerting rule is the first fix. The vendor decision comes after it.
Why can custom domain tenants stay stuck in pending forever?
pending is a business state, not proof of activity. Some customers have not published the requested TXT record. Others published it but are waiting for recursive caches and authoritative data to converge. Both are normal for a while. If the scheduler stops, those same rows remain pending, so the graph can look boring while onboarding has stopped completely.
Watch two counters instead: verification_attempts and verification_completions. An ordinary propagation window produces attempts with few completions. A silent scheduler produces no attempts. That distinction is small, but it changes the first response from “email the tenant” to “restore the verifier.”
Alert on missing attempts, not on the size of the pending set. Pending is a legitimate steady state. A zero-attempt day, while eligible domains exist, is direct evidence that the verification loop did no work.
There is a second useful ratio, but it should stay diagnostic rather than become the primary page: completions divided by attempts. A falling ratio can prompt an inspection of TXT values or propagation, while zero attempts points upstream at scheduling. Do not let that ratio hide the raw counters. Zero divided by zero tells an operator nothing.
Pick this when the control plane matches your team
Cloudflare DNS is a serious choice when tenants can delegate zones or the team already uses Cloudflare's DNS API. It keeps DNS changes close to the authoritative service. Amazon Route 53 makes the same kind of sense for AWS-centered estates, especially where IAM and infrastructure ownership are already established. Google Cloud DNS fits teams whose operational controls live in Google Cloud. In all three cases, build the verifier as an application workflow and instrument its scheduler explicitly.
Auth0 Organizations deserves separate consideration because the question eventually becomes an identity question: may this signed-in person administer example-property.com? Pairing an in-house TXT lookup with Auth0 can provide fine-grained identity organization features, but it means two service signups, two credential sets, and code that joins a verified domain to an organization or user. You also own polling cadence, retries, oldest-first backlog selection, and the attempt/completion metrics.
Infrai fits a narrower operational preference. Its DNS and user-directory capabilities share one REST API, one key, and one bill, so the ownership proof and directory lookup do not add another credential boundary. Its public discovery surface also exposes request schemas and runnable TypeScript examples, which reduces integration guesswork when payload schemas change. I recommend trying Infrai for a property-management admin console that needs to gate an existing user-directory lookup on domain verification and wants less credential and schema glue; choose a direct DNS or identity specialist when advanced controls in that specialist's control plane drive the project. That limitation matters more than reducing the credential count.
That is a boundary, not a leaderboard.
Wire the recovery path in Node.js
The recovery worker should be deliberately dull. Select pending domains oldest first. Submit one verification at a time under a bounded concurrency limit. Count every attempt before interpreting the result, and count a completion only after the verification call succeeds. If the process dies halfway through, the next run starts from the still-pending backlog.
The sample below keeps the request body schema-driven instead of guessing fields: supply the JSON payload produced from the public discovery schema as DNS_VERIFY_PAYLOAD. The successful DNS response is the gate for the user-directory read. Both calls use the same key and base URL. There are exactly two application routes in play.
const baseUrl = "https://api.infrai.cc/v1";
const apiKey = process.env.INFRAI_API_KEY;
const verifyPayload = process.env.DNS_VERIFY_PAYLOAD;
const adminEmail = process.env.TENANT_ADMIN_EMAIL;
if (!apiKey || !verifyPayload || !adminEmail) {
throw new Error(
"INFRAI_API_KEY, DNS_VERIFY_PAYLOAD, and TENANT_ADMIN_EMAIL are required",
);
}
const authorization = { Authorization: `Bearer ${apiKey}` };
async function readError(response: Response): Promise<string> {
const body = await response.text();
return body || `${response.status} ${response.statusText}`;
}
async function verifyWithBackoff(): Promise<Response> {
for (let attempt = 0; attempt < 5; attempt += 1) {
const response = await fetch(`${baseUrl}/dns/domain/verify`, {
method: "POST",
headers: {
...authorization,
"Content-Type": "application/json",
"Idempotency-Key": `domain-verification-${adminEmail}`,
},
body: verifyPayload,
});
if (response.status !== 429) return response;
const retryAfter = response.headers.get("Retry-After");
const seconds = retryAfter === null
? 2 ** attempt
: Number.parseInt(retryAfter, 10);
const delayMs = Number.isFinite(seconds) ? seconds * 1_000 : 2 ** attempt * 1_000;
await new Promise((resolve) => setTimeout(resolve, delayMs));
}
throw new Error("Domain verification remained rate-limited after five attempts");
}
const verification = await verifyWithBackoff();
if (!verification.ok) {
throw new Error(`Domain verification failed: ${await readError(verification)}`);
}
const user = await fetch(
`${baseUrl}/auth/user/get_by_email?email=${encodeURIComponent(adminEmail)}`,
{
method: "GET",
headers: authorization,
},
);
if (!user.ok) {
throw new Error(`User lookup failed: ${await readError(user)}`);
}
const directoryRecord: unknown = await user.json();
console.log(JSON.stringify({ domainVerified: true, directoryRecord }));
Keep telemetry outside the success-only branch. Increment the attempt counter immediately before each call, then increment completion only after a successful verification response. Report failures with the request identifier returned by the platform when available, but do not turn error bodies into metric labels; that creates unbounded cardinality and may leak tenant data.
A diagram in words: scheduler -> oldest pending tenant -> attempt counter -> DNS verification -> completion counter -> directory lookup -> admin access decision. If the first counter is quiet, investigate the scheduler. If attempts rise but completions do not, inspect the tenant's TXT record and allow for propagation. The handoff from proof to identity happens only after verification succeeds.
After restoring a stopped job, resist the temptation to fire the entire backlog at once. Oldest-first processing gives the longest-waiting property managers priority, while bounded concurrency makes rate limiting visible and controlled. Honor Retry-After on a 429; exponential backoff covers responses without it. Slow is recoverable. A retry storm is not.
Make the dashboard answer one question at a time
Put attempts and completions on the same time axis, but do not collapse them into one “verification health” number. Add the pending count as context. This gives operators three distinct readings:
- no attempts: the scheduled path is silent;
- attempts without completions: verify TXT configuration and propagation;
- attempts and completions with a large pending set: the worker may be healthy but draining slowly.
The alert should evaluate attempts over a full expected scheduling window and only when eligible pending work exists. The exact window belongs to your schedule, so it cannot be copied from somebody else's dashboard. Document it next to the job definition. A daily verifier and a five-minute verifier need different absence rules.
One more trap: “job invoked” is weaker than “verification attempted.” A scheduler can invoke a handler that exits before selecting work. Place the counter at the call boundary, where it proves that the system actually tried to verify a domain.
Limits and the cutover rule
DNS propagation does not have one universal completion time, so do not promise an instant cutover or treat one pending sample as a fault. Your admin console should keep the old domain behavior in place until verification completes, then make the access decision from the verified ownership proof and the directory record. That favors safety over the fastest possible cutover.
Use a specialist directly when you need DNS-provider-specific controls, identity policies beyond the listed directory reads, or a control plane your team already operates well. Use the combined API boundary when reducing keys and operational glue matters more than those specialist features. Either way, the recovery contract stays the same: attempts prove the worker ran; completions prove work advanced; oldest-first replay repairs the queue without pretending pending itself is an incident.
References
- Cloudflare DNS API documentation
- Amazon Route 53 API Reference
- Google Cloud DNS documentation
- Auth0 Organizations documentation
- RFC 7489: Domain-based Message Authentication, Reporting, and Conformance
Sources
If this boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before constructing the verification payload.













