DOGFOOD 2026 had one brief for every team: in 72 hours, build a self-hostable hackathon submission and judging portal. Same spec, same anonymized dataset, and a small program, run.py, that pokes your running portal and writes a report you have to commit. No taking your README's word for it.
Mine is called Verdict. This is the story of the four things that taught me the most: one architecture decision, one Docker surprise, one algorithm I nearly shipped, and one line of the spec that looked trivial.
Read the checker before writing code
The first thing I did was read run.py. It runs 7 checks, 3 for tier 1 and 4 for tier 2, and none for tiers 3 and 4. Two of the checks shaped the whole build:
- "Project from fixtures shown" reads the gallery's raw HTML. A client-rendered gallery looks perfect in a browser and fails the check. So the gallery is server-rendered.
- "Judge cannot see peer scores" uses two judges who share 5 projects and a track, so only peer isolation can stop the request. No track rule hides it for you.
The dataset had surprises too. It's described as 40 projects, but there are 41: one is a resubmission of another, same team and same repo, 3 minutes before the deadline. Three judges scored both copies and disagreed with themselves (one gave 2.33 to one copy and 3.67 to the other). Merging them would double-count judges, so the later copy is flagged as a duplicate and excluded.
The decision that shaped everything: Postgres, and rules in the database
My starting plan was the familiar Mongo + Express + Next + Node stack. I swapped Mongo for Postgres before writing a line of code. Two reasons. The rules of a judging platform are invariants, and invariants belong where no code path can skip them. And Mongo transactions need a replica set, which is one more thing to break in a one-command docker compose up.
So the database itself enforces:
- Scores inside the rubric's range. A trigger rejects anything else, whatever the application code does.
- Immutable results. A published results run is a snapshot that triggers refuse to update.
-
An append-only audit log. Each row includes a hash of the previous one, and
UPDATEon the table is blocked.
It paid off within the first 20 minutes. My UNIQUE(event, team name) constraint failed on import: 3 team names in the official data are each shared by different teams, with different members. The constraint was wrong, not the data. Letting the database reject my assumptions early was cheaper than finding out from a judge.
The Docker gotcha: "it works offline"
The portal had to run with no internet: no CDNs, no cloud APIs. Fonts are bundled from npm, and everything ships in the images.
Then, on container startup, npx prisma migrate deploy printed a notice that a new major version of npm was available. To know that, npx was quietly asking the npm registry about updates on every container start. Online, it's a harmless notice. Offline, it's a request waiting to time out on the startup path. The fix is boring: call the binary directly.
./node_modules/.bin/prisma migrate deploy
A cousin of this bit me outside Docker: npm 11 blocks package install scripts by default, so Prisma's engine and esbuild silently didn't install on my machine. Inside the container it was fine, because the Node 22 image ships npm 10. Two environments, two package managers, two different behaviours. Before submitting, I ran the whole stack on a network with no route out at all and checked that the checker still passed 7/7.
The algorithm I almost shipped: z-scores
The spec asks for "a documented way to even out judges who score harshly or generously". The textbook answer is per-judge z-scores: subtract each judge's mean, divide by their standard deviation. It's what I would have shipped.
Instead I simulated it first. I kept the dataset's exact judge–project pairs, generated 300 events with known true quality, harsh and generous judges and rubric-style rounding, then measured how well each method recovered the true ranking (Spearman ρ):
| Method | Mean ρ vs truth | Picks the true winner |
|---|---|---|
| Per-judge z-score | 0.747 | 29% |
| Raw average | 0.788 | 37% |
| Additive model + shrinkage | 0.806 | 40% |
Z-scores were worse than doing nothing. Judges here review 3 or 4 projects each, and a standard deviation estimated from 3 numbers is noise. Dividing by noise amplifies it.
What replaced it is one line of maths:
score = overall mean + project quality + judge leniency + noise
A judge's leniency is learned from the projects they share with other judges, then removed. Shrinkage (ridge regularization) pulls small-sample estimates toward zero, so a judge with three reviews isn't over-corrected. It's about 60 lines of TypeScript solved by coordinate descent, with no maths library. I also ran a control with perfectly fair judges, which prices the insurance: the model costs a little there (0.829 vs 0.851 for raw averages). That's a trade-off I'd rather state than hide.
The finding nobody wants: the judges didn't agree
Then the uncomfortable result. I measured inter-rater reliability (ICC) on the official data: the judges agree with each other no better than chance. Two reviews of the same project differ about as much as reviews of different projects, on every criterion.
No normalization can create signal the scores don't contain. When I added rank uncertainty (resample each score from its noise and re-rank, many times), even the #1 project's likely rank spans #1 to #18, with roughly even odds of finishing top 4.
Most platforms would print a confident podium anyway. Verdict shows a 90% rank range for every project, and its Integrity page warns organizers to treat close ranks as ties. Being honest about uncertainty turned out to be the most important "feature" of the judging engine.
The integrity checks needed the same discipline. Tuned against real data instead of guesses:
- The dataset has only 4 distinct comments, so naive copy-paste detection flagged "Docs are thin." 29 times. Now it ignores anything under 25 characters.
- "Judge opposite to the panel" first fired at a correlation of −0.06, which is noise. The threshold is now −0.3, which leaves the one judge who really is opposite (r = −0.99).
- Later, the head-to-head mode flagged an honest judge who agreed with the panel 33% of the time. Over about ten comparisons, that happens by chance a quarter of the time. A judge is now flagged only when a binomial test says chance can't explain it (p < 0.05).
The line I thought was simple: "assign judges"
"Assign 3 judges to each project, respecting tracks and conflicts." A greedy balancer, I thought.
My first version produced a valid assignment that split the judges into 4 islands that never shared a project. That quietly breaks normalization: you can't compare one island's harsh judge to another island's generous one if no project connects them.
The cause was in the data. One track had 6 projects and exactly 3 eligible judges, so all three had to review all six. The two who also covered other tracks were fully booked before they could act as bridges. A "prefer bridging judges" tie-break wasn't enough. The fix was a repair pass: swap single assignments to a judge in another island, and keep a swap only if the number of islands drops. On the official data it gives one connected graph, every project at 3 reviews, and loads of 2–7 (the dataset's own assignments range from 1 to 11).
My test had asserted that loads spread by at most 3. On this data that's impossible, because the bottleneck track forces 6. I changed the test to assert what the data allows.
I attacked my own app before submitting
Writing the threat model, I checked every claim against the running system instead of the code. That found four real bugs:
-
A forgeable client IP. The proxy setup trusted
X-Forwarded-Forfrom the client, which defeated per-network vote and sign-in limits. - No limit on password guessing.
- Account takeover. Registering with the email of an invited judge who hadn't set a password yet set the password and signed the attacker in.
- The viewer's name in the public embed's HTML, hidden by CSS but still there.
All four were fixed before the freeze, each with a regression test. An integration test now calls every documented API route as every kind of user who should be refused, more than 600 calls, and checks both the refusal and that nothing changed.
What I'd tell myself at hour 0
- Let the database reject your assumptions. Constraints and triggers find data surprises before users do.
- Simulate before you trust a formula. The textbook answer was the wrong answer for 3-review judges.
- Show the uncertainty. A ranking without error bars is a claim you can't defend.
- Verify your docs against the running system. Every bug in the security section came from checking a sentence I had already written.
The final build passes the official checker 7/7, with 416 unit tests and 189 integration tests against a real Postgres. The normalization proof regenerates from one command, deterministically.
Repository: https://github.com/theseriouscoder123/dogfood













