The self-healing I wired in last night didn't work at all this time, and while bringing back missing stocks I reversed my diagnosis twice.
This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).
Last night's fix didn't hold up
Yesterday(new tab) I wired self-healing logic onto the backup GPU path — erase a dead slot, bring it back.
That night's replay test ran 100 out of 100 clean, so I trusted it.
Then on the first real production night, one slot died around 2am, and within about an hour all four slots were dead.
Self-healing fired over 200 times but barely revived anything. One worker got stuck on a single stock for 45 minutes before timing out, and the whole run stalled at less than half its target. The GPU sat idle for hours before I noticed and cleaned it up by hand.
What I asked an AI, and what it said back
I wondered whether an AI advisor's earlier pushback against raising concurrency — from a few days ago — had anything to do with tonight's collapse.
So I asked it directly: given how badly this went, doesn't it owe an apology for that call?
It went back through the logs itself and answered. No apology owed for the concurrency call — tonight's failure mechanism had nothing to do with that setting, and going the other way would have been the same or worse.
What it did owe an apology for was something else: trusting that "erase and revive" alone was enough, and pushing a full server-restart fallback down the priority list. In practice, that fix's success rate was close to zero.
I took the correction and shipped five fixes the same day. Killed the wasteful call pattern that let a single stock eat 45 minutes, and added an escalation path — if a slot dies again shortly after being healed, restart the whole server instead.
I also added a watchdog that restarts if no worker produces output for a while regardless of cause, a fix so a restart doesn't redo already-finished stocks, and an optional evidence log for the healing process.
Eight stocks came back empty the next morning
I resumed the halted run the next morning, and 8 of the 100 target stocks came back empty.
My first guess was an external price API error. Wrong — the code already tolerates that failure gracefully.
My second guess was a bug in the restart logic I'd just fixed. Wrong again — the logs showed that logic never even triggered that morning; there was no situation for it to fire on.
The real cause was much simpler. There's a safeguard that auto-stops GPU work during regular market hours, and I forgot to disable it when manually resuming the run that morning.
So the moment the market opened, all four workers safely stopped a few stocks short each. It was the same mistake I'd made once before.
There was a worse problem underneath. I'd filled in the 8 missing stocks using the last values sitting in the log — but cross-checking again showed those values were from a run several days earlier, not today's. The same log file had multiple days' records stacked on top of each other, and I'd missed that.
So instead of trusting the log, I pulled those 8 stocks out and reran them for real. During that isolated rerun, I also found and fixed a separate bug where the monitoring code was watching an already-finished production output file, wrongly concluded the job had stalled, and force-restarted the server three times.
Two attempts later, all 8 stocks were filled with real rerun results, and the full 100 were confirmed.
Letting it run through market hours this time
That morning I made a call: for this one run, explicitly override the rule that auto-stops GPU work during regular market hours.
The reasoning was simple — there was no real reason it had to finish by the usual cutoff.
So I let it run without the auto-stop, and by afternoon all 100 stocks came out clean and the GPU was released properly.
Two smaller cleanups
One was a false alarm bug in holiday timers. I'd thought I'd fixed it a few days earlier, but the evidence I relied on — "no alert fired on Sunday" — turned out to be meaningless.
Those jobs never even attempt to run on Sundays regardless of the holiday check, so the absence of an alert proved nothing. I re-verified with three separate layers instead: the logic itself, the actual shell script behavior, and the install state.
The other was a documentation fix. A note about a promotion-detection bug had been sitting marked "not yet fixed," when it had actually been fixed and deployed days earlier. Just a stale record I caught and corrected.
What's next
Next weekend (October 3rd) is the decision point on whether to keep this backup GPU path as-is or reorder the hardware priority.
How reliably it holds up over the next few days — after today's collapse and the five fixes that came out of it — is what that decision will hinge on.
!["[260929] R9700 Backup GPU Total Failure - Getting the Cause Wrong Twice"](https://media2.dev.to/dynamic/image/width=1200,height=627,fit=cover,gravity=auto,format=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhdgp7b06eo6k0ihqiu0g.png)



!["[260925] 뉴스 수집 범위 실시간화와 보조 카드 안전장치 배포"](https://media2.dev.to/dynamic/image/width=1200,height=627,fit=cover,gravity=auto,format=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fibmjiil2r7ftj4ahqly1.png)


!["[260922~260927] 안전장치를 겹겹이 쌓고, 전제를 다시 검증한 한 주"](https://media2.dev.to/dynamic/image/width=1200,height=627,fit=cover,gravity=auto,format=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff2wh1c9v3z72p19jqfxn.png)
!["[260928] R9700 Backup GPU - Wiring Self-Healing In and Rechecking the Concurrency Bump"](https://media2.dev.to/dynamic/image/width=1200,height=627,fit=cover,gravity=auto,format=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3nccwp7nogp812vlxm1h.png)