Football/ Ayush Pawar

05 / Transferability Complete

Will his game travel?

What happens to attacking output when a player changes leagues?

Similarity can describe a player.
Transferability asks whether his output survives the move.

  • 428Summer moves
  • 142Locked test moves
  • 5Leagues

Finding

Most output survives a move on average (97% vs comparable stayers), but one league-difficulty ladder decides the direction. A model of the player's own history, age and clubs beats “he'll do what he did” on unseen seasons; similarity adds nothing to it.

02Why it matters

Every summer, clubs pay for output produced somewhere else. Whether that output travels is the question behind every cross-league signing.

Study 4 can say who a player resembles. It can't say what happens after he moves.

How much of a player's output survives a league move — and can we predict it?

03What we did

  1. Player history
  2. Origin league
  3. Transfer
  4. Destination league
  5. Post-transfer output

The benchmarkWe compare movers against — same club, role, output and age — so that ordinary isn't mistaken for a transfer effect.

Output drops after any strong season, move or no move. Comparing movers with similar players who stayed separates the move from that ordinary drop.

Technical detailHow the estimate is built

Understat has no transfer records, so moves were inferred from club changes (including deadline-day moves after early matches). Of 2,917 cross-league moves, the primary population is 428 summer moves (2016/17–2025/26) by forwards and midfielders with ≥900 minutes before and after; every exclusion is listed with its reason.

= output after ÷ output before (xG + xA per 90). Dates of birth were linked from the CC0 Transfermarkt-derived dataset (100% of the 428 moves).

Every prediction uses only information available at the decision date; the final model was scored once on .

04What we found

On average, most of it survives.

Movers keep 97% of their output compared with comparable stayers (95% 93–101%): an average move costs about 3%.

Output kept after a move, vs comparable stayers

xG + xA per 90. 100% = no change relative to similar players who stayed. Line is the 95% interval.

Movers keep 97% (95% interval 93–101%): an average move costs about 3%, and the interval just reaches 100%.

Study 05 · 428 summer moves, 2016/17–2025/26

But the direction decides it.

One explains all 20 directions: moving into the Premier League keeps 82%; moving out of it, 117%.

Premier League hardest, Serie A and La Liga in the middle, Bundesliga and Ligue 1 easiest. Pair-specific effects add nothing (p = 0.91): Serie A → Premier League keeps ~80%, Premier League → Serie A ~114%.

Output kept moving into vs out of each league

vs comparable stayers, xG + xA per 90. 100% = no change.

Moving into the Premier League costs the most and moving out of it gains the most; the Bundesliga and Ligue 1 are the reverse.

Study 05 · 428 moves; per-league 95% intervals in the table below

Age and the destination club matter too.

Younger players keep more — about −2% per year of age (U21 109%, 30+ 93%). The destination club matters as much as the league: weakest third of destination clubs 79%, strongest third 106%.

Players coming off a jump in output drop more (−9% per SD of last-season change). No effect: role (once regression to the mean is removed), consistency over three seasons, profile distinctiveness, penalty duties, loans.

Output kept vs comparable stayers, by age and destination club

xG + xA per 90. 100% = no change.

Under-21s keep 109%, players over 30 93%. Joining a club in the weakest third keeps 79%; the strongest third, 106%.

Study 05 · 428 moves

Can we predict the move?

Each step adds information the naive guess ignores. The final model cuts the error from 0.129 to 0.096 and lifts from 0.58 to 0.76.

Prediction error on 142 locked test moves, step by step

First season after the move, xG + xA per 90. Lower MAE is better; higher R² is better.

Each step adds information and lowers the error: 0.129 → 0.117 → 0.102 → 0.096.

Study 05 · locked test, 2023/24–2025/26

The model was frozen before the test.

  1. Training seasons
  2. Model locked
  3. 2023/24 — 2025/26 opened
  4. 142 unseen moves, scored once

The final model was fixed before the test seasons and evaluated once on 142 unseen moves. It beat the naive baseline by 0.033 xG + xA per 90 (CI 0.018–0.049) and did slightly better on the unseen seasons than on earlier ones — no sign of overfitting.

Similarity had to earn its place. It didn't.

Adding Study 4's similarity — what the most similar destination-league players produce, match quality, distinctiveness, how similar movers fared — changed the error by 0.000. Comparables alone were worse than the player's own history.

Prediction error on 157 league moves

Mean absolute error, xG + xA per 90, first season after the move. Lower is better.

Similar players alone barely beat the naive guess. The player's own history cuts the error by 17%; adding similarity on top moves it by less than 0.001.

Study 05 · critical experiment, transfers with Study 4 profiles

A prediction is a range, not a number.

The chance of keeping ≥75% ranks players well ( 0.80) and beats a same-for-everyone base rate ( 0.179 vs 0.224), and is at the top. But individual forecasts are uncertain: a typical is ±0.2 xG + xA per 90.

Lower thresholds (50%, 60%) were not well calibrated and are not used.

Is the chance of keeping ≥75% honest?

Five groups of about 29 test moves. On the diagonal = says what it means.

Where the model says about 82% and 96%, 86% and 97% of players kept ≥75%. Lower down it is less reliable. AUC 0.80; Brier 0.179 vs 0.224 for a same-for-everyone guess.

Study 05 · locked test, calibration by fifths

Known weakness

Where the model breaks.

The league effect is a fixed amount, not a percentage, so low-output central midfielders moving into the Premier League are under-predicted by about 0.06 xG + xA per 90.

Example: Manuel Ugarte

Manchester United case study, summer 2024 — standing at 1 June 2024 with only what was known then.

xG + xA per 90 · Study 05 case study

Ugarte was bought for defensive work this data cannot see. The model primarily sees attacking output. A percentage-based league effect is the next version.

Read the case study

05What surprised us

  1. What the raw data suggested

    Early on, centre-forwards seemed to drop the most after a move, and a promoted-club effect looked like a data artefact.

  2. What happened after we checked

    Against comparable stayers, centre-forwards keep 96% — mostly regression to the mean, no clear role effect. And the promoted-club effect is real: players joining promoted clubs produce more than the club's rating predicts.

  3. Lesson

    Every raw transfer pattern needs a control group — and a second look.

06What it means

  • To forecast a move, start from the player's own history, age and both clubs, and place both leagues on the difficulty ladder.
  • Use similarity to find candidates and the transfer model to size the risk — never one as the other.
  • Read every forecast as a range. The model knows how unsure it is; a single number hides that.

07Limitations

  • Output only: defending, pace, possession, injuries, fees and contracts are not in the data; defenders and goalkeepers are excluded.
  • The league effect is fixed rather than proportional — the Ugarte weakness. A percentage-based version is to be validated on moves after 2025/26.
  • Individual forecasts are uncertain: a typical 80% range is ±0.2 xG + xA per 90.
  • Observational data: effects are associations within carefully matched comparisons.

SourcesUnderstat · football-data.co.uk · Open-Meteo ERA5 · Wikipedia/Wikidata · OpenStreetMap
Analysis periodMoves 2016/17 — 2025/26 · frozen dataset · Methodology

08The next question

Five studies, one answer: context belongs to the match; level belongs to the player.

Does it still hold on seasons the studies never saw?

Live tracker