Thirty judges. 126 ballots. One judge, jdg_07, scored every criterion of every project a 4.
Standard deviation: exactly 0.
That one judge is why a hackathon portal of mine ended up with a ridge-regression normalizer. This is my write-up for the Hackathon Raptors Dogfood challenge: every team builds the same product, the portal Raptors would run its own events on, in 72 hours. I built Raptors, a self-hosted submission and judging portal (FastAPI, SQLite, server-rendered Jinja), and most of the interesting decisions in it were forced by the data.
TL;DR
- Built Raptors, a self-hosted hackathon submission and judging portal (FastAPI, SQLite, server-rendered Jinja) for the Hackathon Raptors Dogfood challenge.
- A judge who scores everything 4 breaks the textbook per-judge z-score, so I built a two-stage calibration: shrunken scale, then a ridge additive model.
- Calibration cut the spread between judges by 40.7%, but I also show what that does not prove, and how sensitive the result is to my parameter choices.
- I claimed only the tiers the acceptance script can verify: T1 and T2, 7/7 PASS.
Why one flat judge matters: a plain average treats his 4s as real evidence, and on a small event a few judges like him decide the podium. So calibrating judges stopped being a nice-to-have and became the core product requirement. The rest of this post is how I got there, and just as much, what I can and can't claim about it.
Repo: AryanSaxenaa/raptors (MIT). docker compose up --build gives you a seeded portal.
Read the grader before writing a line
The brief ships a spec, a fixture file and an acceptance script, run.py. Before building anything, I read all four closely. INTEL.md in the repo is a forensic reading of all four, written before any implementation code. PLAN.md records each engineering decision in the same shape: what most teams will do, why that is weak, what we do instead.
Reading the grader first paid for itself immediately. These are the traps the obvious build walks into:
The one worth repeating outside this challenge is the permission check. A hidden resource (404) or an empty list (200) is not a denial. The peer-scores endpoint returns a real RFC 9457 peer_scores_denied with a 403 and writes a row to the audit log. The rest are specific to this grader: the gallery renders all 41 projects in arrival order, because the suite only reads page one.
The architecture decision that shaped everything
The suite greps the raw HTTP bytes of the gallery. A client-rendered gallery looks perfect in a browser and fails that check, so the gallery is server-rendered Jinja. That one constraint set the shape of the whole app:
- Public reads are HTML. Every write is JSON through
/api/*, with no second, privileged form-post path. If the UI can do something, curl can do it through the same endpoint. - Auth is one opaque token (
df_sessioncookie or Bearer), because the suite attaches a cookie and never logs in. - One process, one SQLite file in WAL mode. No Redis, no Postgres. I wrote down when I would move to Postgres: measured need, not because it sounds production-grade.
The Docker gotcha: a port that answers on an empty database
If the server binds its socket before it seeds the database, there is a window where the app answers requests with no data in it, which is how a gallery check fails on a perfectly healthy container.
So bootstrap (schema and fixtures) runs before the socket binds, and /api/health doesn't mean "process alive". It reports seed counts and requires them to match the fixtures and the schema_meta.version:
{"ok":true,"schema_version":"3","schema_ok":true,
"counts":{"projects":41,"judges":30,"scores":126},
"fixture_counts":{"projects":41,"scores":126}}
The scoring algorithm that doesn't survive the data
The obvious approach is a mean per project. On these fixtures, it is vulnerable to large judge effects:
- The highest judge mean is 4.22. The lowest observed judge mean is 2.00, from
jdg_01, who submitted only one ballot, so it says little about how harsh that judge really is. - Loads run from 1 ballot to 11 per judge.
- Eight projects got 2 reviews and four got 5.
A plain mean treats a generous 4 and a harsh 4 as the same evidence. The textbook fix is a per-judge z-score, and it breaks twice. It divides by zero on jdg_07. For a single-ballot judge the spread is undefined, so an implementation has to choose a fallback. If that fallback is a zero spread, the normalized score becomes 0.0 and the observation contributes no discriminative information. It is also poorly suited to this assignment structure, because judges were restricted by track and so did not necessarily score comparable samples.
What shipped is two stages in pure Python with no solver:
-
Shrunken scale. Each judge's spread is blended with the pooled spread, weighted by their ballot count:
σ̃ = (nσ + kσ_pool) / (n + k), withk = 5. Atk = 5, a judge needs more than five ballots before their own spread gets more weight than the pooled spread. As long as the pooled spread is nonzero,σ̃is nonzero, so a flat judge likejdg_07cannot cause a divide-by-zero (the code falls back to a scale factor of 1 if every ballot in an event were identical). Each judge's scores are then rescaled around a shrunken mean. Forjdg_07that gives every project they scored the same value, so their ballots say nothing about which of those projects is better. That is a choice I made: a judge whose scores contain no within-judge variation doesn't get to discriminate between projects. It doesn't mean their ballots have no effect on the ranking. -
Ridge additive model.
score = μ + project effect + judge bias, fitted by alternating least squares on the cells that exist. The penalty is the same for every judge, but the data's pull against it is not. With 11 ballots, the data can overcome the judge penalty. With one ballot, the penalty pulls that judge's bias strongly toward zero, because one observation cannot separate "generous" from "good project".
On the published fixtures, the spread between judges falls from 0.4127 to 0.2449, a 40.7% reduction:
But is the new ranking better, or just different?
This is the question that matters, and a smaller spread alone doesn't answer it. 29 of the 40 ranked projects change position after calibration. Small Meadow and Flat Meadow each drop 9 places and Flat Relay rises 6:
So I fit the same estimator on a labelled synthetic panel: the same (judge, project) cells as the fixtures, scores generated from planted project effects and judge biases. Because the truth is known, I can score the recovery directly:
| Check on the planted panel | Result |
|---|---|
| Correlation, recovered vs true project effect | 0.872 |
| Correlation, recovered vs true judge bias | 0.963 |
| Kendall τ vs true ranking, raw means | 0.649 |
| Kendall τ vs true ranking, calibrated | 0.671 |
The gain on that one planted panel is modest (τ from 0.649 to 0.671), but one panel is one random draw. So I repeated the experiment over 200 random seeds, with the same missing cells as the fixtures. Treat this as a simulation check that the estimator recovers what was planted, not as empirical validation on real judging data:
| Synthetic panel, 200 seeds | Mean τ, raw mean | Mean τ, calibrated | Calibrated better in |
|---|---|---|---|
| Matches the model's assumptions | 0.540 | 0.698 | 199 of 200 seeds |
| Judges apply different scales to project effects | 0.549 | 0.693 | 198 of 200 |
| Project effects include a nonlinear quadratic term | 0.512 | 0.663 | 199 of 200 |
The single panel above understated the gain. Across the three scenarios, the mean τ improvement ranges from 0.144 to 0.158. These are three separate scenario results, not one pooled estimate. I still treat this carefully, because the planted panels are variations on the same additive idea the estimator assumes.
What this does not prove
- Calibration is not ground truth. No one knows the true best project of a real event, and the planted-data test only shows the method recovers effects I planted.
- The synthetic panels are friendly to the estimator. The judge-scale and quadratic variants are mild misspecifications, not an adversarial test.
- The 40.7% figure measures spread, not accuracy. It is a drop in how spread out the judges' averages are, and says nothing about whether the ranking is more accurate.
- Identifiability is limited. Assignment is restricted by track and isn't random, so judge and project effects are only weakly separated for judges with few ballots. That is why the portal flags any project with fewer than 3 reviews as low confidence, and prints a rough standard error next to each score (the model's overall residual spread divided by the square root of the project's review count). It is a warning signal, not a confidence interval.
-
k = 5and the ridge penalties are design choices, not discovered constants. They come from reasoning, not a search. Atk = 5, a judge needs more than five ballots before their own spread gets more weight than the pooled spread (the median load is 3). The judge penalty is larger than the project penalty because a judge with few ballots is confounded with whatever they happened to be assigned.
Here is how much the shrinkage constant and the ridge penalties matter on the real fixtures (τ against the default calibrated ranking):
| Setting | τ vs default | Top 3 unchanged | Top 10 overlap |
|---|---|---|---|
| shrinkage k = 1 | 0.938 | yes | 10 of 10 |
| k = 3 | 0.972 | yes | 10 of 10 |
| k = 5 (default) | 1.000 | yes | 10 of 10 |
| k = 10 | 0.982 | yes | 10 of 10 |
| k = 20 | 0.959 | yes | 9 of 10 |
| judge penalty 0.5 | 0.903 | yes | 10 of 10 |
| judge penalty 1 | 0.946 | yes | 10 of 10 |
| judge penalty 3 (default) | 1.000 | yes | 10 of 10 |
| judge penalty 12 | 0.954 | yes | 10 of 10 |
| project penalty 0.25 | 0.990 | yes | 10 of 10 |
| project penalty 1 (default) | 1.000 | yes | 10 of 10 |
| project penalty 3 | 0.977 | yes | 10 of 10 |
Across the parameter values I tested, the top 3 never changes, and the top 10 changes by one project at most (at k = 20 the overlap is 9 of 10). That is evidence about these values, not a proof that the ranking is invariant to the parameters. The middle of the table does move a little. A lighter judge penalty actually scores better on the synthetic panels (mean τ 0.728 at 0.5, against 0.698 at 3), so the default isn't the optimum among the tested synthetic penalty values. I kept the heavier penalty because of the confounding argument above, and I'd rather say so than pretend the number was optimised.
What it changes for an organizer
On the real fixtures, calibration swaps the top two and puts a different project on the podium.
-
prj_37replacesprj_10in the top 3, and the top-10 membership is the same. - Eight projects move five or more places, and 29 of 40 move at all.
- Dropping
jdg_07's three ballots entirely changes the order of 18 projects, by up to 10 places (τ 0.941 against the full ranking), with the same top 3. This also changes how many reviews some projects have, so it is not a pure measure of that judge's influence, but it does show that a flat judge still moves ranks through the ridge stage even though they can't tell their projects apart. - With
prj_41included in the ranking, the duplicate would rank 8th, a top-10 spot. With the duplicate excluded, as shipped, the originalprj_07ranks 27th.
Reproduce the proof with python -m dogfood.normalize --proof --fixtures fixtures.json, and the tables above with the script at the end of this post. The repo was frozen for judging before I wrote the sensitivity analysis, so the script isn't in it.
The part of the spec I thought was simple: "judges can't see each other's scores"
The deadline
"Closed events refuse submissions" sounds like one if. But if you validate the body first, a malformed payload gets a 422 instead of "closed". The deadline check now runs before the JSON is parsed, and tests/test_deadline.py throws malformed JSON at a closed event to prove it.
The UI isn't the control
The role matrix is 5 actors by 5 capabilities, 25 cells. The brief's checker probes three. All 25 are enforced in one ROLE_MATRIX in security.py and walked over HTTP in tests/test_role_matrix.py. Every denial is a problem document with a machine-readable code, and it writes an authorization.denied row to a hash-chained, append-only audit log. Hiding a button isn't isolation. curl is:
$ curl -i -H "Cookie: df_session=<judge_b>" "localhost:8080/api/judge/scores?judge=jdg_01"
HTTP/1.1 403 Forbidden
content-type: application/problem+json
{"code":"peer_scores_denied","status":403,
"detail":"scores belonging to judge jdg_01 are not visible to you"}
The duplicate that took 4 of 126 review slots
The fixtures contain a duplicate submission: prj_41, a copy of prj_07 (Dry Harbour). The duplicate submission accounts for 4 of the 126 ballots, and the two copies together account for 9. Three judges (jdg_19, jdg_21 and jdg_26) reviewed both copies.
The rule is that the duplicate is flagged, excluded from the ranking, and its ballots still count for calibration. Dropping them would bias the judges who saw both copies. The assignment algorithm also skips known duplicates when topping up reviews, so scarce judge attention isn't spent on something that won't be ranked.
What I refused to claim
The brief lets you claim a tier only if you can prove it, and says that overclaiming is the one thing that costs points. run.py verifies T1 and T2, so .dogfood.toml claims T1 and T2, and the report closes with claimed T1 T2, verified T1 T2. Quadratic community voting, signed records, an embeddable gallery, bulk import/export and the full API (T3 and T4) are built, but I didn't claim them because the checker can't verify them. I also skipped pairwise judging on purpose: the brief says to pick one hard bonus and nail it, so I picked normalization.
How I worked
I used Cursor for the implementation, under rules I set: the spec is the contract, run.py is the grader, the reading of the checks goes in INTEL.md and PLAN.md before any application code, and everything is verified over HTTP. I claimed T1 and T2 only, and the report reads 7/7 PASS. About 80 pytest cases pass (25 of them are the role matrix), plus a black-box smoke suite of about 190 checks.
Audits and tests caught the real late bugs, and each fix has a test: the open and email voting modes returned 403 because the ballot endpoints still demanded a judge permission, an organizer could publish results while voting was still open, a participant could move a project onto another event's track, and the default gallery mixed two events.
What I took from it
The hard part of a judging portal isn't accepting scores. It's deciding what those scores are allowed to mean when the data is sparse, biased, duplicated, probed by a script, and graded by software. Every decision I'm proud of in this project is about that question, and the ones I'm least sure of, such as k = 5, I've tried to show rather than hide.
Try it
git clone https://github.com/AryanSaxenaa/raptors
cd raptors
docker compose up --build
# in another terminal
python run.py .dogfood.toml
Walkthrough
A silent 80-second walkthrough of the portal, with background music:
Appendix: the sensitivity script
Save it as sensitivity.py in the root of a clone of the repo and run python sensitivity.py. It only reads fixtures.json and imports dogfood.normalize.
sensitivity.py
"""Sensitivity of the calibrated ranking to k_shrink and the ridge penalties.
Save this file in the root of a clone of the repo and run: python sensitivity.py
"""
import json, random, statistics, sys
from pathlib import Path
ROOT = Path(__file__).resolve().parent
sys.path.insert(0, str(ROOT / 'src'))
from dogfood import normalize as N
fx=json.load(open(ROOT / 'fixtures.json'))
scores=fx['scores']
keys=sorted({k for s in scores for k in s['criteria']}); w={k:1.0 for k in keys}
obs=[N.Observation(s['judge'],s['project'],N.weighted_value(s['criteria'],w)) for s in scores]
pids=[p['id'] for p in fx['projects']]
EXCL={'prj_41'}
def rank(res,method):
r=[p for p in res.projects if not p.excluded and p.raw is not None]
r.sort(key=lambda p:(-(p.score(method)),p.project_id)); return [p.project_id for p in r]
base=N.compute(obs,all_project_ids=pids,excluded_project_ids=EXCL)
raw=rank(base,'raw'); cal=rank(base,'additive_ridge')
print('ranked',len(cal))
print('top3 raw',raw[:3],'cal',cal[:3],'same set top3',set(raw[:3])==set(cal[:3]))
for k in (3,5,10):
print('top',k,'overlap',len(set(raw[:k])&set(cal[:k])),'of',k)
pos=lambda l:{x:i for i,x in enumerate(l)}
pr,pc=pos(raw),pos(cal)
print('moved>=5 places',sum(abs(pr[x]-pc[x])>=5 for x in cal),'moved>=1',sum(pr[x]!=pc[x] for x in cal))
# duplicate check: does excluding vs including change top 10
inc=N.compute(obs,all_project_ids=pids,excluded_project_ids=())
ci=rank(inc,'additive_ridge'); print('dup rank if included', ci.index('prj_41')+1 if 'prj_41' in ci else None, 'prj_07 rank', cal.index('prj_07')+1)
def tau(a,b):
ids=[x for x in a if x in set(b)]
return N._kendall_tau(N._ranks([-pos(a)[x] for x in ids],ids=ids),N._ranks([-pos(b)[x] for x in ids],ids=ids))
print('--- sensitivity on real fixtures: tau vs default calibrated ranking')
for k in (1,3,5,10,20):
r=rank(N.compute(obs,all_project_ids=pids,excluded_project_ids=EXCL,k_shrink=k),'additive_ridge'); print('k',k,round(tau(cal,r),3),'top3',r[:3]==cal[:3],'top10 overlap',len(set(r[:10])&set(cal[:10])))
for lj in (0.5,1,3,6,12):
r=rank(N.compute(obs,all_project_ids=pids,excluded_project_ids=EXCL,lambda_judge=lj),'additive_ridge'); print('lambda_judge',lj,round(tau(cal,r),3),'top3',r[:3]==cal[:3],'top10 overlap',len(set(r[:10])&set(cal[:10])))
for lp in (0.25,1,3):
r=rank(N.compute(obs,all_project_ids=pids,excluded_project_ids=EXCL,lambda_project=lp),'additive_ridge'); print('lambda_project',lp,round(tau(cal,r),3),'top3',r[:3]==cal[:3],'top10 overlap',len(set(r[:10])&set(cal[:10])))
print('tau raw vs default',round(tau(cal,raw),3))
# synthetic multi-seed, well-specified and misspecified
pj=sorted({s['project'] for s in scores}); jd=sorted({s['judge'] for s in scores})
def synth(seed,mode,k=5,lj=3,lp=1):
rng=random.Random(seed)
P={p:rng.gauss(0,.55) for p in pj}; B={j:rng.gauss(0,.8) for j in jd}
S={j:(rng.uniform(.5,1.6) if mode=='scale' else 1) for j in jd} # per-judge scale
o=[]
for s in scores:
j,p=s['judge'],s['project']
if mode=='scale': v=3.5+B[j]+S[j]*P[p]+rng.gauss(0,.3)
elif mode=='skew': v=3.5+B[j]+P[p]+rng.gauss(0,.3)+ (0.5*P[p]**2)
else: v=3.5+B[j]+P[p]+rng.gauss(0,.3)
o.append(N.Observation(j,p,min(5,max(1,v))))
res=N.compute(o,all_project_ids=pj,k_shrink=k,lambda_judge=lj,lambda_project=lp)
sc=[x for x in res.projects if x.raw is not None]; ids=[x.project_id for x in sc]
tp=N._ranks([P[i] for i in ids],ids=ids)
return N._kendall_tau(tp,N._ranks([float(x.raw) for x in sc],ids=ids)), N._kendall_tau(tp,N._ranks([float(x.additive_ridge) for x in sc],ids=ids))
for mode in ('well-specified','scale','skew'):
a=[synth(s,mode) for s in range(200)]
print(mode,'raw tau mean',round(statistics.fmean(x[0] for x in a),3),'cal tau mean',round(statistics.fmean(x[1] for x in a),3),'cal>raw in',sum(x[1]>x[0] for x in a),'/200 seeds')
print('--- synthetic (well-specified, 200 seeds) by params')
for k,lj in [(1,3),(5,3),(10,3),(5,0.5),(5,12)]:
a=[synth(s,'well-specified',k=k,lj=lj) for s in range(200)]
print('k',k,'lj',lj,'cal tau',round(statistics.fmean(x[1] for x in a),3),'raw',round(statistics.fmean(x[0] for x in a),3))
Built for DOGFOOD 2026 by Hackathon Raptors.























