The charter
Most hackathons now run the same way. Wire up a few APIs, put a chat box on top, demo it before it breaks. Gaussathon is for people who work on problems where the data is thin or broken and creative problem reframing often helps. You are scored on the number you produce and on the reasoning behind it.
- 01
- No language models. Everything can't be a one-shot NLP problem.
- 02
- No agents or copilots. We are protecting the one weekend a year where the modelling and thinking is yours.
- 03
- Method over horsepower. A gradient-boosted kitchen sink will not save you.
- 04
- Show the derivation. Every entry carries a short note on the model, its assumptions and where it breaks. Report your uncertainty.
- 05
- Everything opens at the close. Code and notes go public when scoring ends, as a citable record of how the same problem got attacked five ways.
Three tiers of judgement
Scored in order rather than as a weighted sum. Tier 1 decides who reaches the finals, tier 2 separates the finalists, tier 3 breaks ties.
- Tier 1
- Performance. One metric, announced in advance, scored automatically on a held-out set you never see. The public leaderboard runs on a subsample and the private one decides. Where the target is a distribution we use a proper scoring rule.
- Tier 2
- Elegance. Whether your model says something about how the data was generated, judged from your note by a panel. Parameters have to be earned, and assumptions have to be stated and tested.
- Tier 3
- Complexity. Structure you took on and made work, such as hierarchy, latent variables, measurement error or censoring. Ambition without diagnostics scores nothing.
The panel and the full rubric are published with the problems, before the clock starts. Judges recuse themselves from teams they have co-authored with.
The problems
Released at kickoff. Real datasets from working scientists and industry partners, each one nasty in its own way.
- Track A
- Estimation under scarcity. A few hundred rows, thirty candidate covariates, and one coefficient somebody will make a decision with.
- Track B
- Signal in a hostile process. Heavy tailed, irregularly sampled, and the instrument changed halfway through the record.
- Track C
- Many rows, few labels. Millions of measurements and a few hundred labels, each of which cost an expert an hour. Show what the unlabelled mass bought you.
- Track D
- Structure without supervision. Twenty-eight instruments across equities, FX, rates and metals, updating hourly, with no labels. Learn a representation and defend it on probes you do not get to pick, against a production embedding.
- Wildcard
- Unseen until day two. One problem nobody gets in advance, including most of the judges. Same three tiers, weighted for the shorter clock.
Track D data comes from Heston Labs, whose Latent API streams the series and the 128-dimensional embedding used as the baseline. Gaussathon's organiser founded Heston Labs and does not judge that track.
Have a dataset that belongs in this room? Propose a problem.
Format
Two days, teams of one to four. Times are indicative until the venue is confirmed.
- Day 1, 09:00
- Problems, data dictionaries, baselines and the scoring harness all go live at once.
- Day 1, 10:00
- Open modelling. Judges and invited statisticians are around to consult, and asking them a good question is encouraged.
- Day 1, 19:00
- Chalk talks. Optional, whiteboard only, and somebody explains an estimator they love.
- Day 2, 09:00
- The wildcard opens for anyone who wants it.
- Day 2, 16:00
- Submissions close and the private leaderboard is revealed.
- Day 2, 17:00
- Shortlisted teams defend for eight minutes and take questions. Prizes, then the archive goes public.
- Teams
- 1 to 4 people
- Cost
- Free to attend
- Compute
- Identical, capped, provided
- Languages
- R, Python, Julia, Stan, anything
- Places
- Limited, applications reviewed
Questions
Is the no-LLM rule enforceable?
Not completely, and we are not pretending otherwise. Submitted environments are inspected, every shortlisted team defends its derivation in front of a panel that will ask why you chose that link function, and the whole archive goes public afterwards. Anyone who cheats at a statistics hackathon has already lost the part worth winning.
Can I use scikit-learn, XGBoost, Stan, PyMC, brms?
Yes. The ban covers language models and coding agents, not established libraries. Neural networks are allowed where they suit the problem, though on these datasets they rarely do, and tiers 2 and 3 will ask you to justify one.
Do I need a PhD?
No. We care about how you think about uncertainty, not your credentials. Applications are reviewed because places are limited, not to screen for titles.
What do I submit?
Predictions in the specified format, a code bundle that runs end to end inside the provided environment, and a technical note of at most four pages. The note is what tiers 2 and 3 are scored from.
Can I come without a team?
Yes, and many do. There is a team-forming session before the problems drop, and solo entries are competitive.
Who owns the work?
You do. Publishing at the close is a condition of entry, under a permissive licence you choose, but the work and any paper that follows are yours.
Come and estimate something hard
Registration is by application and places are limited. It takes two minutes: who you are, what you have modelled, and which track you want.