Every team that touches URLs eventually has the same argument. Should decoding happen in the browser before a form submits, or should it wait for the server? Should the proxy decode before forwarding, or pass the raw string through and let the application decide? The choice looks small, but it shapes bug reports, security posture, and how easily new engineers can reason about request data.
This guide is for engineers who already know what percent-encoding is and want a framework for deciding where decoding belongs in a real system. It is not a tutorial on the syntax itself — for that, the Lizely guide to decoding a URL in your browser without writing code is a clean walkthrough. The goal here is to make the architectural call easier.
The Decision Is About Boundaries, Not Syntax
Percent-encoding is a transport representation, not application semantics. A URL component like %2F is the same character as / once decoded, but the wire format encodes it differently to disambiguate parsing. RFC 3986 defines the reserved and unreserved sets and how each character must travel (RFC 3986, Section 2). Every framework agrees on that grammar. The disagreement starts at the seam between layers.
Three questions anchor the decision:
-
What does the downstream consumer expect? A router matching against
/users/:idtypically wants the decoded value bound to:id. A cache key built from the raw path needs the encoded form so that/users/a%2Fband/users/a/bdon't collide. - Who owns canonicalization? If two layers decode independently, you risk double-decoding, which silently corrupts data. One layer must be the single source of truth.
-
What happens when decoding fails? Malformed escapes like
%XYshould fail loudly at a boundary you control, not inside business logic where the stack trace points at the wrong culprit.
Answer those, and most debates resolve themselves.
Server-Side Decoding: The Default That's Usually Right
Most frameworks decode path segments and query parameters automatically before route handlers run. Express with express.urlencoded, Spring with its @RequestParam binding, Django via request.GET, and ASP.NET Core out of the box — they all parse and decode as part of the request pipeline. MDN's reference page on the URL API is a useful cross-check when you are unsure which layer is responsible.
The reasons this default wins most of the time:
-
Routing requires decoding. Path-based routing cannot tell
%2Ffrom/if it does not decode first, so it must. - Audit and logging live on the server. If your access logs are the canonical record of what happened, decoded values are easier to grep and to render in a UI.
-
Security scanners expect decoded input. WAFs and IDS rules are written against the decoded form, because attackers don't actually send
%3Cscript%3Efor the encoded form to matter.
The common failure mode is implicit server decoding combined with explicit client decoding. A frontend that decodes a parameter before adding it to a query string, then a backend that decodes again, produces a single decode where the developer expected two. The fix is usually to remove the client-side call, not to add another server-side one.
Client-Side Decoding: When the Browser Owns the Truth
There are legitimate cases for decoding in the browser. They share a pattern: the application needs to read or display the value before sending it anywhere.
-
Showing the user what they actually typed. A "share this link" preview should display spaces, not
%20. Decoding on render avoids that. -
Validating user input before submission. A form that rejects characters reserved by RFC 3986 should check the decoded value, not the encoded one — otherwise an end user typing
?into a search box sees a confusing error. -
Comparing two URLs as semantic equality.
encodeURIComponentis not a hash. If you need to know whetherhttps://x/a%2Fbandhttps://x/a/bpoint at the same resource, decode before comparing. The WHATWG URL standard (HTML Living Standard, URL chapter) treats the parsed interface as the source of truth.
The trap on the client is double-decoding user-visible text that was already decoded by the framework when it was first stored. The first decode is part of the data model. Subsequent decodes are bugs.
The Proxy and CDN Layer: A Frequent Source of Confusion
Reverse proxies, CDNs, and API gateways sit in front of the application and frequently offer a "decode URL" toggle. The toggle is dangerous because it usually applies only to path normalization, not to query string semantics. Engineers enable it expecting routing to fix itself, then discover that a request for /files/report%2F2024.pdf arrives at the origin as /files/report/2024.pdf and breaks the lookup.
A practical rule for the proxy layer:
- Decode only for the purposes of routing and cache keys.
- Pass the original encoded form to the origin unless the application has explicitly opted in to receiving decoded paths.
- Never decode twice. If the proxy decodes for routing, the origin must not decode again.
If the cache key includes the path, build it from the encoded form. Decoded-path cache keys cause the well-known "cache poisoning by trailing slash" class of bug where two semantically different URLs share an entry.
A Production Checklist for Teams
This is the checklist I use when a URL-decoding question lands in a code review or a postmortem. Paste it into your team wiki and adapt the wording.
- Identify the single owner of decoding for each URL component (path, query, fragment).
- Confirm that the owner's behavior matches the framework default — if you are overriding
urlencoded, document why. - Verify that cache keys use the same encoded form the proxy uses for routing.
- Add a request-log line that records both the raw path and the decoded path on errors only.
- Write one unit test per route that submits a value containing a reserved character (
/,?,#,+,&). - Reject malformed percent-escapes at the boundary, not deep in business logic.
- Forbid manual decoding in frontend code unless it is paired with a comment naming the component and the reason.
Items 5 and 6 catch most regressions. The rest prevent the next one.
What "Correct" Looks Like in Practice
A concrete worked example helps two teams agree on what "right" means.
Suppose the route is GET /search and the frontend wants to send a query q=c++ & more (note the spaces and ampersand). The naive frontend builds ?q=c++ & more, which the server parses as two parameters because & is a delimiter. The naive fix is to encode the whole query string, which then encodes the = and ? and breaks the request.
The correct sequence is:
- Frontend calls
encodeURIComponenton each value, not on the whole string. - Frontend joins encoded values with
=and&to form a syntactically valid query. - Browser sends
?q=c%2B%2B%20%26%20more. - Server decodes each parameter value once.
- Server validates and returns results.
If the frontend instead calls decodeURIComponent on what the user typed before encoding, it has done a no-op on already-readable text and introduced a place where a malformed escape could throw. Don't add that step. The user typed readable text; the encoder's job is to make it safe for transport, starting from the unencoded form.
The same shape applies to path components, hash fragments (which are never sent to the server), and form-encoded POST payloads. The principle is invariant: encode at the producer, decode at the consumer, and pick exactly one of each.
When the Rules Bend
A few situations justify breaking the default:
- Legacy origins that already decoded in business logic. Wrapping them in a proxy that does not decode lets the existing code keep working. Document the boundary in the proxy config.
- Signed URLs with signatures over the encoded path. The signature is over the bytes the client signed, which is usually the encoded form. Don't decode before verification.
- Internationalized domain names. Punycode is its own encoding. Treat it as a separate concern from percent-encoding, and apply the standard conversion in one place.
Each of these is a place where the simple rule "decode once at the boundary" is correct, but the boundary is non-obvious. The mitigation is the same: name the boundary in code, write a test that pins its behavior, and revisit when the next refactor touches that area.
Frequently asked questions
Should I ever decode inside a React or Vue component?
Only to display human-readable text. Never to build another URL. If you find yourself decoding to compose a string you'll re-encode, stop — pass the original value through and let the next layer encode.
What's the difference between decodeURIComponent and URLSearchParams?
decodeURIComponent handles a single string. URLSearchParams parses plus a full query string and decodes each value, returning an iterable map-like object. Use URLSearchParams whenever you have a query string; it avoids manual split('&') mistakes and handles the +-as-space convention correctly. For a single component, decodeURIComponent is fine.
My server logs show %20 but the UI shows spaces. Which is the bug?
Neither, on its own. Decide which is the source of truth for your team. Most teams pick the encoded form for logs (because it is what was on the wire) and the decoded form for error pages (because humans read them). The bug to look for is mismatched expectations — a search tool that grep-decodes before matching, or a log viewer that renders spaces and breaks correlation with the proxy.
How do I write a regression test for a decoding bug?
Pick one reserved character per test. Submit it through the production code path, assert that the handler receives the decoded value and that the response is correct. Then submit a malformed escape (%ZZ) and assert the server returns a 400, not a 500. Two tests, one route, ten minutes to write — and they prevent the entire "silent corruption" class of decoding bug from coming back.
This article was drafted with AI assistance and reviewed for technical accuracy before publishing.







