When a sort key vanishes entirely, a sort function quietly hands back the original order instead
This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).
Every morning, my automated trading system pulls the largest stocks by market capitalization from two domestic markets and uses that list as the day's trading universe.
One morning last week, that list changed almost entirely overnight. And one stock that had no business being on it slipped in — and actually got bought.
What happened
I found out by accident. Someone asked, "why did we buy this stock?" — and only then did I realize something was off.
Looking at the logs, most of the top names in both markets had been swapped out that morning. There was no reason for market-cap rankings to shift that drastically in a single day.
There was no error in the logs. The refresh job had finished with its usual "completed successfully" status.
What I thought first, and why it was wrong
My first suspicion was an unusual market move. My second was a bug in the universe-construction logic.
Both were wrong. I only found the real problem after tracing all the way back to where the data was actually fetched.
The real cause
That morning, the open-source library the system uses to fetch market data returned the market-cap column as entirely missing values (NaN) for every row.
The logic that picks the top N stocks by market cap sorts on that column. When the sort key is entirely missing, the sort function doesn't raise an error.
Instead, it hands back whatever order the raw data happened to already be in — in this case, something close to alphabetical/code order. The system took that result and treated it as "the top N by market cap," no questions asked.
The end result was a universe full of stocks that had nothing to do with actual market capitalization, and one of them made it all the way through to a trading signal and an actual buy on the live account. Fortunately it wasn't a directional loss — just a small transaction cost from a trade that never should have had a signal behind it.
The fix
To make sure the same failure couldn't slip through silently again, I added validation in several independent layers.
The first layer validates at ingestion. When fetching market data, the system now checks what fraction of stocks have a valid market-cap value, and rejects the entire source if too many are missing. If the live feed gets rejected, it falls back to the most recent clean cached snapshot instead.
The second layer is an anomaly tripwire. If the universe's membership changes by an abnormal amount in a single day, that alone is treated as a red flag — the update gets held back and the previous state is kept. This exists to catch contamination that somehow slipped past the first layer for any reason.
The third layer is a final sanity check right before the result is actually used, after it has already passed the first two.
The recovery itself was a matter of rolling back to the last known-good state. But partway through, another process that hadn't been restarted yet re-contaminated the universe using the old logic — which meant the recovery order had to change to "restart every process first, then restore state," not the other way around.
The general lesson
When a sort key column goes entirely missing, sorting doesn't fail — it produces a plausible-looking fake result. No error means the logs look clean. If a ranking or ordering step depends on external data, "the sort finished" and "the sort finished meaningfully" are two different claims.
An external data source can pass its schema check and still be empty. A column that exists, with the right type, but is entirely null won't get caught by a typical schema validator. You need a separate check on the actual distribution or valid-value ratio of the data.
One layer of input validation isn't always enough. This recurred even after input validation was in place, which is why a separate anomaly layer — watching for "the output changed by way too much" — was added on top. When the cause of a contamination is hard to pin down in advance, monitoring the symptom (an abnormal magnitude of change) tends to be more reliable than trying to enumerate every possible cause.
Recovery order is itself something to verify. While rolling back a contaminated state, any other process still running on the old logic can re-contaminate it. If multiple processes share the same state, recovery has to bundle "roll back the state" together with "make sure every consumer of that state is on the new code" — not treat them as separate steps.


