Football/ Ayush Pawar

Research · 01—05

What actually belongs to the player?

Five studies investigating how much football performance belongs to the player — and how much belongs to the context.

  • 2015/16 — 2024/25Frozen dataset
  • 5Leagues
  • 18,011Matches
  • 8,348Players

The research spine

One question. Five investigations.

Each study builds on the one before. The first three remove context from the numbers; the last two use the cleaned numbers to compare players and follow them across leagues.

  1. 01 / Environment 2015/16 — 2024/25 Weather freeze pending

    Does weather change performance?

    Mostly no. Crowds do.

    Weather barely changes how much players produce: with all five leagues, rain and humidity effects sit within about ±2%, and wind is small (about −1% per +10 km/h). Crowds clearly do: empty stadiums roughly halved home advantage, and it came back with the fans.

    Read the study
  2. 02 / Home advantage 2015/16 — 2024/25 Complete

    How much does playing at home matter?

    A lot — and the same for everyone.

    Everyone does better at home by about the same amount: +27% xG per minute for the same player, in the same team and season. Not one of 4,665 players has a reliable personal home edge.

    Read the study
  3. 03 / Opposition 2015/16 — 2024/25 Complete

    What happens when the opponent gets stronger?

    Everyone drops — by about the same amount.

    Player attacking output falls substantially as opponent strength increases — but we found no reliable evidence that some individuals consistently resist that effect.

    Read the study
  4. 04 / Similarity Profiles to 2025/26 Complete

    Can we find players who actually play alike?

    Yes — within limits.

    Context-adjusted profiles find sensible matches across leagues — Bruno → Serie A: Samardžić, Dybala, Chukwueze — but nobody matches his chance creation. Similarity describes a player; it is not a forecast.

    Read the study
  5. 05 / Transferability Moves 2016/17 — 2025/26 Complete

    Can we predict what happens when the player moves?

    Partly — from his own history, not from who he resembles.

    Most output survives a move on average (97% vs comparable stayers), but one league-difficulty ladder decides the direction. A model of the player's own history, age and clubs beats “he'll do what he did” on unseen seasons; similarity adds nothing to it.

    Read the study

Synthesis

How the studies fit together

Study 1 asked whether the environment moves performance — weather barely, crowds clearly. Study 2 took the crowd effect to the individual: home advantage is real, large and shared by everyone. Study 3 did the same for opposition: large, shared, with no individual “big-game” trait.

Together they gave clean, context-adjusted player numbers. Study 4 used those numbers to find similar players across leagues, and found the limit: similarity describes, it does not forecast. Study 5 tested forecasting directly on real transfers and built a validated, leakage-free model of how much output survives a move — in which similarity had to earn its place and did not.

What we learned

Five lessons that run through everything

  1. Context moves the numbers, and it moves everyone the same.

    Home, crowd and opposition effects are large, but they belong to the match and the league, not to individual players or clubs. Every “special” player we looked for — home specialists, big-game players, crowd-sensitive players — disappeared once noise was separated from signal.

  2. Most apparent individual differences are small samples.

    A player's home/away split is only 5–8% signal, and his strong/weak-opponent split is mostly noise; raw rankings of these traits are in action (Study 2: the correlation between players' 2015–19 and 2022–25 home edges is 0.018).

  3. A player's overall level is the best predictor of almost everything.

    His overall opponent-adjusted output predicts his output against elite teams better than his record against elite teams (Study 3); his own pre-move history predicts his post-move output better than similar players do (Studies 4 and 5).

  4. Similarity is a discovery tool, not a forecasting tool.

    Comparable players help find candidates; they add nothing to a forecast once the player's own history is known (Study 5, critical experiment).

  5. Leakage discipline matters.

    Understat's match “forecast” turned out to be post-match; it was never used. Opponent strength is built only from pre-match odds; every Study 5 prediction uses only information available at the decision date, and the final model was scored once on .

Corrections

Where we corrected ourselves

The interesting result isn't always the first one. These are the readings we changed along the way.

  • Understat's win/draw/loss “forecast” was at first described as pre-match; it is post-match (r = 0.96 with the match's xG difference). It was never used in any model.
  • Study 2's “stable part” magnitudes are upper bounds: the cross-season covariance estimator is biased upward even with no true persistence (found in Study 3). Conclusions unchanged.
  • Study 5: a promoted-club term was first read as an imputation artefact; it is a real effect (players joining promoted clubs produce more than the club's rating predicts).
  • Study 5: the early “centre-forwards drop most” reading was mostly regression to the mean (96% vs stayers, no clear role effect).

Open

Still open

  • Study 1's hand-written weather conclusion and freeze.
  • Study 4's user-curated “known similar pairs” check.
  • Study 5's percentage-based league effect, to be validated on moves after 2025/26.

SourcesUnderstat · football-data.co.uk · Open-Meteo ERA5 · Wikipedia/Wikidata · OpenStreetMap
Analysis period2015/16 — 2024/25 · frozen dataset · Methodology