Football/ Ayush Pawar

Methodology

How do we know?

A transparent account of the data, adjustments, models, validation and limitations behind the research.

Terms with a dotted underline open a definition · Full glossary →

  1. Data
  2. Context
  3. Adjustment
  4. Model
  5. Validation
  6. Limitations

01 · Data

One dataset, frozen.

Every player in every match of the Premier League, La Liga, Serie A, the Bundesliga and Ligue 1 from 2015/16 to 2024/25 — the frozen study data that all five studies share.

leagues
5
ten seasons
2015/16–2024/25
matches
18,011
with usable player data
18,008
teams
160
players
8,348
player-matches
525,328
shots
~452,000

Per player-match: minutes, goals, assists, shots, , , , , , cards, granular position, home or away, opponent and kickoff. Per shot: location, xG, situation, body part, last action and the assisting player.

Three matches have no usable player data and are excluded. Live seasons are added weekly without altering the frozen dataset: 2025/26 and 2026/27 are estimated the same way and shown beside the frozen results on the live tracker, never pooled into them.

02 · Data sources

Where each piece comes from

  • Understat

    Player-match and shot data: xG, xA, shots, key passes, xGChain, xGBuildup, positions, shot locations and situations.

    Used in: Every study

  • football-data.co.uk

    Pre-match Pinnacle odds for the frozen seasons; market-average closing odds for the live seasons.

    Used in: Opponent strength (Studies 3–5, live tracker); Study 2's market check

  • fixturedownload.com · openfootball

    Kickoff times, reconciled across up to four sources.

    Used in: Weather at the kickoff hour (Study 1); kickoff-hour controls

  • Wikipedia / Wikidata · OpenStreetMap

    Stadium locations (geocoded with Nominatim).

    Used in: Weather at the stadium (Study 1)

  • Open-Meteo (ERA5)

    Hourly reanalysis weather at each stadium and kickoff hour.

    Used in: Study 1

  • Crowd-restriction calendars

    A dated, sourced calendar of crowd status per league, from news reports.

    Used in: Behind-closed-doors comparisons (Studies 1–2)

  • Transfermarkt-derived dataset

    Dates of birth, from the CC0 dcaribou/transfermarkt-datasets.

    Used in: Age in the transfer model (Study 5)

03 · Context adjustment

A player's raw output isn't necessarily his own signal.

The same striker produces more at home, less against a strong side, more in some leagues than others. The research's core move is to take that context out before asking what belongs to the player.

  1. Player
  2. Team
  3. Season
  4. Venue
  5. Opponent
  6. League environment
  7. Context-adjusted output

Compare each player with himself (Studies 1–3)

Why this?
Comparing different players mixes ability, team quality and luck. Comparing one player with himself — same club, same season, home against away or strong against weak opponents — makes those cancel out.
Technical implementation
on per-minute output with (Study 1's weather models use player × season), plus calendar and kickoff-hour terms (league × season × month, league × kickoff hour); standard errors clustered by match.
What it controls for
Who the player is, his club and season, the calendar, kickoff time — and, depending on the study, venue and opponent strength.
What it doesn't solve
It is observational: effects are associations within matched comparisons. Game state (trailing against strong teams) and squad rotation are part of what is measured, not removed.

Adjusted profiles (Study 4)

Why this?
To compare players across leagues, a profile has to describe the player, not his schedule.
Technical implementation
One model per role and metric, not stacked corrections: PPML with a player effect, opponent rating, own-team rating, home and league × season. Each match is then re-weighted to an average opponent, half at home, in an average league-season — .
What it controls for
Opponent strength, venue and league environment.
What it doesn't solve
Own-team strength is estimated but deliberately not applied: players tend to join stronger teams at their best, so the term absorbs form, and applying it made profiles markedly less persistent (movers' xGChain r 0.40 → 0.27).

04 · Opponent strength

Measured before kick-off, never after.

How strong was the opponent? The answer has to come from information available before the match — otherwise the match's own result leaks into the measure of how hard it was.

  1. Pre-match information
  2. Opponent rating
  3. Player performance

0 of 155,282

Replacing every odds price, xG value and score after a random cutoff changed 0 of 155,282 pre-cutoff ratings ( checks across 5 cutoffs × 3 ratings).

A sequential market rating

Why this?
Betting odds published before kick-off are the sharpest public estimate of team strength — and they cannot know the result.
Technical implementation
From pre-match Pinnacle odds: ln(p_home / p_away) = home + r_home − r_away, updated match by match (learning rate 0.3, carry-over 0.9). Each match only sees matches that kicked off at least two hours earlier. Promoted teams start at the mean of last season's relegated teams; the first five matches of 2015/16 are burn-in. Effects are reported per +1 of the rating (1 SD = 0.77 log-odds).
What it controls for
Checked against an xG rating, Elo on results, opening odds and points per game: opponent ratings correlate 0.91–0.95, and the opposition effect is −11.8% to −13.8% per SD on every definition.
What it doesn't solve
Odds price in team news, so a rating reflects the expected line-up, not necessarily the actual one (opening odds give the same answer: −13.7% vs −13.8%).

Unseen-season test (2019/20–2024/25): the market rating's Brier score is 0.582 against 0.577 for the match's own closing odds and 0.652 with no information — about 93% of the way to the market's own forecast.

05 · Statistical models

Counts, per minute, with honest uncertainty

PPML on per-minute output

Why this?
Goals, shots and xG are counts that are often zero, and players play different minutes. A Poisson-type model handles both and reads naturally as percentage changes.
Technical implementation
with minutes as exposure and the fixed effects above; 95% from match-clustered standard errors.
What it controls for
Everything in the fixed effects, plus the specific context terms (venue, opponent rating, weather, crowd status).
What it doesn't solve
It estimates average effects. Whether an individual player differs from the average is a separate, much harder question — next block.

Individual effects: signal vs noise

Why this?
Some players will always look like home specialists or big-game players by chance. Each individual estimate has to be weighed by how much real information it carries.
Technical implementation
Calibrated standard errors (shuffles within each player's own matches), DerSimonian–Laird between-player spread and empirical-Bayes ; persistence measured as the excess over shuffled data ( in action).
What it controls for
Small samples and luck: an estimate built on few minutes is pulled toward the average.
What it doesn't solve
Shrinkage can hide a real but rare individual trait — the studies report that none is detectable, not that none can exist.

06 · Multiple testing

Test thousands of players and some will look special by luck.

When you test thousands of player-level effects, some will look significant by chance. The research uses and / tests to reduce that risk.

  • Home specialists (Study 2)0 of 4,665 players have a reliable personal home edge once noise is accounted for.
  • Big-game players (Study 3)0 of 23,022 player × metric opposition estimates survive the correction.
  • Individual resistance to opposition (Study 3)After calibration, a shuffle placebo finds 3.5–4.6% of players with p < 0.05 — what chance alone produces.
  • Weather (Study 1)Weather shuffled within league × season × month 1,000 times; the placebo flags 4.5–6.9% of cells per test (nominal 5%).
  • Crowds (Studies 1–2)Closed-doors dates shifted 2–4 years into full-stadium seasons: 14 of 15 tests null.

Every headline effect states which correction it survived: Benjamini–Hochberg on cluster and permutation p-values.

07 · Similarity method

Who plays like him?

  1. 9 metrics
  2. Context adjustment
  3. Role normalisation
  4. Cosine similarity
  5. Ranked player profiles

The profile

Why this?
A profile should describe how a player plays, not who takes the penalties or which league he is in.
Technical implementation
Nine metrics: xA, key passes, xGChain, xGBuildup, , shots per 90, and shot style — , and . Context-adjusted, then within one of seven data-derived , equal weights.
What it controls for
Opponents, venue, league environment and role.
What it doesn't solve
Output only — no dribbling, pace, defending, passing volume or age.

The match

Why this?
Two profiles that point the same way are similar even if one player is busier than the other.
Technical implementation
(Euclidean, correlation and Mahalanobis as checks), reported as a : closer than X% of the role pool. A candidate is robust only if it holds across windows and algorithms. Validated without anyone's opinion by : three-season profiles find the same player in the top 9% of the destination league's role pool after a move.
What it controls for
The arbitrariness of any one window or distance measure.
What it doesn't solve
It cannot say whether a player would succeed after a move.

Similarity ≠ prediction

Similarity describes profile resemblance. It is not a transfer-performance prediction. Study 5 tested it directly: adding Study 4's similarity to the transfer forecast changed nothing — similarity finds candidates, it does not forecast them.

08 · Transfer model

What happens when he moves?

  1. Player history
  2. Age
  3. Both clubs
  4. Both leagues
  5. Expected post-transfer output

Model A

Why this?
A player's output after a move depends mostly on what he has done before — but a single good season overstates it, and the destination matters.
Technical implementation
: the player's three-season level of xG + xA per 90, age, minutes, the strength of both clubs, for both leagues and a promoted-club term — learned from earlier moves (428 in total for the calculator; fitted before the locked test for validation). Prediction ranges come from the model's own out-of-sample errors.
What it controls for
Regression to the mean (a three-season level instead of the last season), ageing, and destination difficulty: into the Premier League players keep 82% on average.
What it doesn't solve
League effects are a fixed amount, so low-output midfielders moving into the Premier League are under-predicted — see the known failure below.

The comparable-stayer benchmark

Why this?
Raw before / after comparisons exaggerate the drop: players move after good seasons, and good seasons regress anyway.
Technical implementation
is also measured against — players at their own club, same role, same pre-season output and age — which removes regression to the mean and normal ageing. Movers keep 97% of output against stayers (CI 93–101%).
What it controls for
Regression to the mean and age.
What it doesn't solve
Who chooses to move is not random; the benchmark matches on output and age, not on everything.

09 · Validation

We don't test the model on the data it learned from.

  1. Training · 286 moves
  2. Model locked
  3. 2023/24–2025/26
  4. 142 unseen moves
0.096vs 0.129 naive
0.76vs 0.58 naive
0.80keeps ≥ 75%
0.179vs 0.224 base rate

These are metrics from the locked transfer-model test, not general measures of the entire football research system.

Model A was fixed before the test seasons were opened, fitted on 286 moves (destinations 2016/17–2022/23), and scored once on 142 moves in 2023/24–2025/26. Its error was lower on the unseen seasons than in the earlier folds (0.110) — no sign of overfitting — and its covered 87% of outcomes (conservative). Only the ≥ 75% retention probability is shown on the site: lower thresholds were not .

Elsewhere: Study 4 is validated by self-retrieval on real league moves; Study 1's weather effects by permutation placebos; Studies 2 and 3 are re-tested every week on seasons they never saw, on the live tracker, which keeps the frozen and live results separate.

10 · Limitations

What the data cannot see

The project is primarily about attacking output. That isn't an apology — it is the honest edge of what this data can measure.

  • Defending
  • Passing
  • Carrying
  • Possession
  • Injuries
  • Transfer fees
  • Contracts
  • Tactical fit
  • Cup matches
  • Pace

Not in the data: defensive actions and pressing, passing volume and completion, progressive carries, possession, injuries, salaries, fees and contracts. Cup matches aren't in the data either, so fixture congestion is only partly visible. Goalkeepers are excluded throughout, and defenders from the transfer model.

The closed-doors period is mostly one season, so per-player crowd estimates are noisy, and there are no attendance figures. All effects are associations within carefully matched comparisons, not experiments.

11 · Known model failure

Where the transfer model goes wrong

The transfer model under-predicts some low-output central midfielders moving into the Premier League. Its league effect is a fixed amount, so it wipes out most of a low-output player's forecast: on the locked test, central and defensive midfielders moving into the Premier League were under-predicted by about 0.06 xG + xA per 90.

The research's next version is a percentage-based (log-scale) league effect, to be validated on moves after 2025/26.

See the Ugarte case

12 · Corrections & clarifications

Recorded, not silently edited

Corrections to earlier statements are recorded in each study's findings, not quietly changed.

  1. · Study 1 · 2 · 3

    Understat's win / draw / loss “forecast” was at first described as a pre-match probability. It is post-match (r = 0.96 with the match's xG difference) and was never used in any model.

  2. · Study 2

    The “stable part” magnitudes (e.g. ±7.7% for xG) are upper bounds, not estimates: the cross-season covariance is biased upward even with no true persistence. Conclusions unchanged; Study 3 reports the excess over shuffled data instead.

  3. · Study 5

    The promoted-club effect is real: with a better imputation of promoted clubs' strength the term stays at +0.078 (CI +0.024 to +0.133) — players joining a promoted club genuinely produce more than the club's rating predicts.

  4. · Study 5

    Centre-forwards' “clearest drop” after a move (raw 92%) was mostly regression from their high pre-season level: against comparable stayers it is 96% (CI 91–102%), and role has no clear effect.

13 · How this site uses it

A presentation layer, not a re-run

The site never re-runs the research in your browser. Every number is read from validated outputs exported by the research pipeline; the Similarity Explorer and Transfer Calculator show results precomputed by the studies' own Python engines, checked against them at build time. Each weekly rebuild checks the frozen numbers against a fixed manifest and stops if any of them has moved.

The research project is still open in a few places: Study 1's hand-written weather conclusion and freeze; Study 4's user-curated “known similar pairs” check; and Study 5's percentage-based league effect.