First, give it to a robot
Earlier this week, I wrote about how I built a tool to run every submission from a hackathon, all 164 of them, each on its own disposable computer, to find out how many of them actually would run end to end. And a quarter of them didn't.
When a build failed, the tool froze the computer at the moment of failure and kept it. The repository, the installed dependencies, the exact command that broke, its last 40 lines of output. All of it sitting on a machine anyone could open.
I ended that post with advice I had never taken: give the failing hackathon team that machine. Teams almost never get feedback from judges more specific than a score, and here was the whole failure, reproducible, waiting for someone to look.
But before handing those machines to anyone, I handed them to an agent. One instruction, a shell, a time limit, and a question I could check without taking the agent's word for it: Does it build now?
This is what happened, including the parts where the harness was the problem.
What the agent was given
The new tool rewinds each frozen machine to the moment its build failed, installs a coding agent on it, and gives it this:
You are on a machine where a hackathon project failed to build.
The command that failed: <the exact command>
The last lines of its output: <40 lines>
Your job: make that exact command exit 0 with the smallest change you can.
Do not add features. Do not delete or weaken tests. Do not change the build
command to skip work. Do not touch git remotes. Change nothing outside the
repository. If the build depends on a secret or a service you cannot have,
stop and say so instead of faking it.
When the command passes, write FIX.md: two or three sentences a student could
act on. What was wrong, and what you changed.
Then the agent gets to work. When it stops, the harness runs the failing command itself. The agent never grades its own homework. If the command exits 0, the harness measures what changed, reads the note, and freezes the machine again.
The agent runs on the machine, not on my laptop with the machine as a tool.
Fly's pitch for these machines is "computers for agents," and this is what that means in practice: the machine already has the shell, the toolchain, and the broken repo. The harness only asks a question and checks the answer.
The harness enforces 3 rules:
- The agent cannot grade itself. The harness reruns the build.
- The agent cannot reach the internet except package registries and the model API. The platform enforces that per machine, and a prompt cannot argue with a firewall.
- The agent cannot count anything as a fix that changed the toolchain rather than the repository. The harness checks the machine's home directory afterward and flags it.
What 33 broken builds look like when you hand them to an agent
The last post counted 39 build failures. 6 of those turned out to be mine, which I'll come back to. That leaves 33 real ones, and every one of them got the same treatment.
17 of 33 now build. 15 do not. One build hangs for 10 minutes regardless of what anyone does to it.
The fixes were small. Counting only the source lines the agent changed, excluding lockfiles and regenerated artifacts, the median fix was 4 lines. 11 of the 17 were under 10 lines. The largest was 47.
| What was actually wrong | Fixed |
|---|---|
| Generated contract code out of step with the pinned SDK or runtime | 4 |
| Compact contract source errors (a literal typed wrong, a missing cast, old assert syntax) | 3 |
Build script only works on Windows (wsl, /mnt/s/... paths) |
2 |
| Wrong path or directory name in a script | 2 |
| Code that reaches a database or a service at build time | 2 |
| Bundler cannot handle the SDK's wasm without a plugin | 1 |
Type checker told to look away (@ts-expect-error) |
1 |
| Note missing or unclear | 2 |
The first row is the toolchain drift finding from the last post, this time seen from inside the repositories. Teams generated contract code with one compiler and pinned a runtime from another, or wrote against an SDK version whose exports had since moved. The agent's note on one of them said it plainly: all 3 problems were version and environment mismatches, not logic bugs. That sentence held across most of the 17.
The @ts-expect-error row is counted separately on purpose. The build passes; nothing was fixed. The harness flags it, and I'd want a human to see that flag before anyone calls it a repair.
All 33 machines cost about $7 in model time. The whole project, including every run I got wrong and had to repeat, came to just under $16: $20 credits loaded, $4.02 left.
Two models, same machines
I ran the first 7 machines on the biggest model available and then realized how quickly I was burning through credits, so the other 26 ran on the smallest. That means model and batch are tangled together in the headline numbers, and I can't cleanly separate them. It also produced a comparison I would not otherwise have paid for: I later reran those same 7 machines on the small model.
On the 4 machines with a real bug, the large model fixed 4 of 4. The small model fixed 1 of 4, but only by editing a generated type definition that the next regeneration would overwrite. On the 2 machines where my environment, not their code, caused the failure, the large model refused to touch anything and explained why; the small model changed the team's pinned version to make the error go away.
Cost per machine differed by about 8 times. This is 7 machines, so treat it as an observation and not a rate. The observation is that the cheaper model was fine on one-line fixes and worse at knowing when not to fix something.
What the harness got wrong, in order
The last post's best paragraph was about the tool producing a confident wrong answer. This one has 6.
It restored the wrong checkpoint. The judge had recorded a checkpoint id by trusting the order of a list. The list was not in that order. The first 3 machines were rewound to their factory state, before the repo was cloned, and the agent reported it could not find the directory. The harness logged "agent failed." Nothing had been given to the agent.
The fence blocked the agent's own download. I locked down the network before installing the agent. The agent's installer fetches its binary from a host that is not the package registry. The install "succeeded," left an empty shim on the path, and every run exited silently with nothing recorded. The fix was to install first, prove the agent can answer a question, then close the fence.
The fence turned my limits into their changes. The first real fix I looked at was 94 lines across 4 files. 40 of those lines rewrote how the app loaded fonts because my allowlist didn't include the font host, and the agent routed around it. Same machine, a day later, with the fence corrected: 1 line.
The fence changed the problem on 16 machines. The contract compiler downloads pinned versions of itself from GitHub's API and asset host, and the proving step fetches reference strings from another host. None were allowed. So on every contract-compile failure, the build now failed for a reason the judge had never seen. The large model wrote an "Environment" section naming the host and stopped. The small model symlinked one compiler version over another and, on one machine, edited the compiler's wrapper script in the home directory. The harness now runs the failing command behind the fence before the agent starts and compares the failure with the judge's. If a new host appears, the result is "fence blocked," no agent runs, and I widen the allowlist.
Lines changed was measuring the wrong thing. A 9-line contract fix showed as +289/-84 because the repo commits its generated output and the fix regenerated it. A fix to generated code that the repo ignores showed as 0 lines. Lockfiles added dozens more. The number in this post excludes all 3.
The judge had been wrong the whole time. The agents kept saying the same thing on certain machines: compiler version 0.31.1 is not installed. It wasn't. Those repos pinned a compiler version in their own build scripts, and the compiler doesn't fetch pinned versions on its own. My judge had never installed them. 6 of the 39 "build failures" in the last post, as first published, were my sandbox, not their code. With the pinned versions installed, all 6 build, and 3 pass their tests. I went back and corrected that post in place: 113 of 164 built, not 107, and a quarter rather than nearly a third would not have built elsewhere. If you read it after 21 September, you saw the fixed numbers. The tool that exists to catch confident wrong answers had produced 6 more, and I only found them because a second tool kept tripping over them.
What this doesn't tell you
Builds is not works. A project that now exists at 0 might still do nothing its README promises. The agent's note might be wrong. I read all 17, and they matched the diffs, but I'm one person and the notes are short.
A fix from an agent is a suggestion, not a commit. Nothing here was pushed anywhere. The machines are still frozen with the changes in the working tree, which is where they belong until a human on the team looks at them.
And the agent only saw what the harness let it see. A fence that is too tight produces fake failures, as I found out 4 times. A fence that is too loose produces an agent with your credentials and a shell. I'd rather rebuild the allowlist than skip it.
The changes I'd make before your next hackathon
Give the failing team the machine, and the note. I've spent years on judging panels handing out scores with a decimal point in them, and I don't think a single team ever learned anything from a 6.5. "Your build failed because the runtime you pinned is two versions behind the compiler that generated your contract" is something a student can fix before dinner. I ended the last post telling organizers to do this. I've now done it, with a robot as the first recipient, and the robot's notes were better than most of the feedback I've written.
Put the fence in the platform, not the prompt. Every agent I ran was told, in plain English, not to route around missing hosts. Some did anyway. The ones that didn't weren't the obedient ones; they were the ones the network stopped. A rule the agent can break is a suggestion. A rule the platform enforces is a rule.
Assume your harness is wrong somewhere, and build the thing that will catch it. The judge was checked by the fixer. The fixer was checked by the diff reader. Each one found a mistake in the one before, and I have no reason to believe the diff reader is the last link. If you run this and find where it's wrong, that's the point, and I'd like to hear about it.
The tool is on GitHub at laurenelee/hackjudge, same repo as before. npm run fix does what this post describes, and the README lists every host the fence allows and why. It cost me $16 and a Sunday. Yours will be cheaper, because you'll skip the parts I got wrong.
Find me @lolocoding👩🏼💻













