Football/ Ayush Pawar

Glossary · 50 terms

Football, translated.

The metrics, statistical concepts and modelling terms used throughout the research — explained without assuming you already know the jargon.

Every entry has a simple and a technical definition. Technical definitions use the research's own methods — where a term has a more specific meaning here than in a textbook, the research's meaning wins. How the research works →

Football metrics

xG Expected goals (xG)

Football metrics

Also: expected goals

How likely a shot was to become a goal, based on thousands of similar shots.

Simple
Every shot gets a score between 0 and 1. A tap-in from a yard out might be 0.8; a hopeful strike from 35 yards might be 0.02. Add them up and you get how many goals a player 'should' have scored from the chances he got.
Technical
Understat's shot-level probability of scoring given location, situation, body part and the action before the shot; summed per player-match.
Why we use it
Goals are rare and noisy. xG measures the quality and quantity of chances, which is far more stable from match to match, so it lets small effects (like playing at home) show up.
In the research
At home, the same player produces about 27% more xG per minute than away (Study 2).
#

xA Expected assists (xA)

Football metrics

Also: expected assists

The xG of the shots a player set up with his passes.

Simple
If you put a teammate through on goal and he misses, you still created a great chance. xA gives you credit for the chance, not the finish.
Technical
Sum of the xG of shots directly assisted by the player's pass, regardless of whether the shot was scored.
Why we use it
It measures chance creation without depending on a teammate's finishing — the core of a creative player's job.
In the research
Bruno Fernandes's xA per 90 in 2022/23 was 0.467 raw and 0.460 after adjusting for opponents — the 99.6th percentile (Study 3).
#

npxG Non-penalty xG

Football metrics

Also: non-penalty xg, non penalty xg, npxg

xG with penalties removed.

Simple
Penalties are worth a lot of xG but say little about how a player plays — whoever takes them gets the boost. Removing them makes players comparable.
Technical
xG excluding penalty attempts, so penalty-taking duties don't inflate a profile.
Why we use it
The similarity profile should describe a playing style, not who is on penalties.
In the research
Non-penalty xG is one of the nine metrics in each Study 4 similarity profile.
#

Shots / 90 Shots per 90 minutes

Football metrics

Also: shots per 90, shot volume

How many shots a player takes per 90 minutes played.

Simple
Volume, not quality: a player who shoots often from anywhere scores high here even if his chances are poor.
Technical
Shot count per player-match from Understat, expressed per 90 minutes; one of the nine Study 4 profile metrics and an outcome in Studies 1–3.
Why we use it
Volume and quality behave differently — pairing shots with xG per shot separates 'shoots a lot' from 'gets good chances'.
In the research
Per +1 SD of opponent strength, a player's shots per 90 fall 10.5% (Study 3).
#

Key passes

Football metrics

Also: chance created, chances created

Passes that lead directly to a shot.

Simple
Any pass that sets up a shot, whether the shot is good or bad.
Technical
Count of passes immediately followed by a teammate's shot (Understat definition).
Why we use it
It measures how often a player creates, alongside xA, which measures how good those chances were.
In the research
Per +1 SD of opponent strength, key passes fall 11.4% (Study 3).
#

xGChain

Football metrics

Also: xg chain

The xG of every attack a player was involved in.

Simple
If you were part of a move that ended in a shot, you share the credit — even if you just played a simple pass early on.
Technical
Total xG of every possession that ends in a shot in which the player was involved, credited in full to each player in the chain.
Why we use it
It captures involvement in dangerous attacks, not just the last two touches.
In the research
xGChain is used as an alternative target in Study 5's robustness checks.
#

xGBuildup

Football metrics

Also: xg buildup, xg build-up, build-up

Like xGChain, but ignoring the shot and the final pass.

Simple
Credit for helping build attacks before the final ball.
Technical
xGChain excluding possessions where the player took the shot or made the key pass — isolates build-up contribution.
Why we use it
It distinguishes deep playmakers from players who only appear at the end of moves.
In the research
xGBuildup is part of the 'involvement' group in the Study 4 profile.
#

Per 90 Per 90 minutes

Football metrics

Also: per 90, p90, rate

A stat scaled to a full match, so players with different minutes compare fairly.

Simple
A sub who plays 30 minutes and a starter who plays 90 can't be compared on totals. Per 90 puts them on the same footing.
Technical
Total ÷ minutes × 90. The studies model per-minute rates directly with minutes as exposure.
Why we use it
Every comparison in the series is about rates, not totals.
In the research
Zirkzee was predicted 0.46 xG + xA per 90 at United; he produced 0.47 (case study).
#

xG / shot xG per shot

Football metrics

Also: shot quality, npxg per shot

The average quality of a player's shots.

Simple
Two strikers can both take three shots a game. One takes them from six yards, the other from thirty. xG per shot tells them apart.
Technical
Total xG divided by shots: a shot-quality measure, independent of how many shots he takes.
Why we use it
Volume and quality are different styles. The similarity profile needs both to tell a poacher from a long-range shooter.
In the research
One of the nine metrics in the Study 4 similarity profile, alongside shots per 90 and non-penalty xG (Study 4).
#

Headed share Headed-shot share

Football metrics

Also: headers, headed shots

The share of a player's shots taken with his head.

Simple
Tells you how a player gets his shots — a target man heads far more of them than a winger.
Technical
Headed shots ÷ all shots, derived from Understat's shot-level body-part field; one of the three shot-style metrics in the Study 4 profile.
Why we use it
Two players with the same xG can get it in completely different ways; shot style keeps the similarity profile about playing style.
In the research
Headed share, in-box share and non-penalty xG per shot make up the shot-style part of the nine-metric Study 4 profile.
#

In-box share In-box shot share

Football metrics

Also: inside the box, penalty area, shot location

The share of a player's shots taken from inside the penalty box.

Simple
Low means a player shoots a lot from distance; high means he gets into the box to shoot.
Technical
Shots from inside the box ÷ all shots, derived from Understat shot locations; a Study 4 shot-style metric.
Why we use it
It captures where a player shoots from, which raw shot or xG totals hide.
In the research
Bruno Fernandes's in-box shot share is in the 3rd percentile of his role: a high volume of shots from distance (Study 4).
#

Football context

Home advantage

Football context

Also: home edge, home specialist, venue

How much more a player produces at home than away.

Simple
Compare a player with himself: same club, same season, home games versus away games.
Technical
Home ÷ away ratio of per-minute output for the same player, team and season, from a PPML model with fixed effects.
Why we use it
It separates what belongs to the venue (and the crowd) from what belongs to the player.
In the research
+27% xG per minute at home; behind closed doors that fell to about +8–14% (Study 2).
#

Behind closed doors

Football context

Also: empty stadiums, no fans, covid, crowds

Matches played without fans during COVID-19 restrictions.

Simple
A natural experiment: the same teams and players, suddenly without a crowd.
Technical
Dated per-league crowd-restriction calendars; mostly one season, so per-player crowd estimates are noisy.
Why we use it
It shows how much of home advantage is the crowd itself.
In the research
Without fans the xG home edge fell from +28% to +12%, then came back to +25% (Study 2).
#

Opponent strength Opponent strength (market rating)

Football context

Also: opposition, market rating, big-game players, strong opponents

How good the opponent was, measured from pre-match betting odds.

Simple
Bookmakers' odds before kick-off are a very good summary of how strong a team is. We turn them into a rating that never peeks at results after the match.
Technical
A sequential rating built only from pre-match Pinnacle odds; each match uses only earlier matches. Checked against an xG rating, Elo, opening odds and points per game.
Why we use it
Using only pre-match information means the rating can't leak the result it is trying to explain.
In the research
Per +1 SD of opponent strength, a player's xG falls about 14% (Study 3).
#

Context-adjusted output

Football context

Also: adjusted, context adjustment, raw vs adjusted

A player's output with the effect of opponents, venue and league environment taken out.

Simple
What the player would have produced against average opposition, in an average setting.
Technical
Per-match output re-weighted to an average opponent, half at home and an average league-season environment, from a PPML context model (Study 4); opponent-adjusted per-90s in Study 3.
Why we use it
Raw output carries its context with it. Adjusting lets players from different leagues and schedules be compared — though for forecasting a move it added little.
In the research
Bruno Fernandes's xA per 90 in 2022/23: 0.467 raw, 0.460 adjusted for opponents (Study 3).
#

Statistics

SD Standard deviation (SD)

Statistics

Also: standard deviation

A standard unit for 'how far from average' something is.

Simple
Roughly, one SD up in opponent strength is the jump from an average side to a clearly strong one.
Technical
Standard deviation; effects are reported per +1 SD of the opponent rating, and profiles are z-scored within role.
Why we use it
It puts very different scales (odds, xA, shots) on a common footing.
In the research
Bruno's chance creation is 3.6 SD above his role average; Serie A's best is 1.9 SD (Study 4).
#

z-score

Statistics

Also: standardised, standardized, z score

How many SDs a number sits above or below the average.

Simple
0 means exactly average for the position; +2 means far above it.
Technical
(value − role mean) ÷ role SD; computed within role so a full-back is compared with full-backs.
Why we use it
It lets nine different metrics be combined into one profile without one dominating.
In the research
Study 4 profiles are z-scored within each of seven data-derived roles.
#

95% interval Confidence interval (95%)

Statistics

Also: confidence interval, ci, uncertainty

The range the true value plausibly sits in, given the data.

Simple
A single number hides uncertainty. The interval shows how sure we are: narrow means confident, wide means 'early days'.
Technical
95% confidence interval from match-clustered standard errors (or bootstrap where stated).
Why we use it
Live-season figures early in a season have wide intervals — the interval tells you not to over-read them.
In the research
Early in a live season the home-edge interval is wide; it narrows as each weekend's matches are added (live tracker).
#

Correlation (r)

Statistics

Also: r, pearson, spearman

How closely two numbers move together, from −1 to +1; 0 means no relationship.

Simple
If players who were good at something last season are also good at it this season, the correlation is high. If knowing last season tells you nothing, it is near zero.
Technical
Pearson correlation unless stated (Spearman for rank comparisons).
Why we use it
It is the simplest test of whether a 'trait' is real: something that belongs to a player should persist from one period to the next.
In the research
Past home edge vs later home edge: r = 0.018 across 595 players (Study 2) — past home edge says nothing about the future.
#

Percentile

Statistics

Also: rank, ranking

Where a value ranks in a group, from 0 (lowest) to 100 (highest).

Simple
99th percentile means higher than 99% of the comparison group.
Technical
Percentile rank within the stated comparison pool (usually the player's role, adjusted profiles).
Why we use it
It makes very different metrics readable on one scale, always relative to a stated group.
In the research
Bruno Fernandes's adjusted key passes are in the 100th percentile of attacking mids / wide forwards (Study 4).
#

Signal vs noise

Statistics

Also: noise, luck, reliability

How much of a difference is real, and how much is chance.

Simple
If you flip a coin 20 times, some 'players' get 14 heads. That doesn't make them better at flipping. Most individual splits in football are like that.
Technical
Share of observed between-player variance that is true variance, from calibrated standard errors and DerSimonian–Laird estimates.
Why we use it
Without separating the two, you crown players for luck.
In the research
A player's home/away split is only 5–8% signal (Study 2).
#

MAE Mean absolute error (MAE)

Statistics

Also: mean absolute error, prediction error, error

The average size of a prediction's miss.

Simple
If MAE is 0.10, predictions are off by about 0.10 xG + xA per 90 on average.
Technical
Mean |actual − predicted| in xG + xA per 90 on the test moves.
Why we use it
It's the headline measure for comparing transfer models.
In the research
Adding similarity changed the error by 0.000 (Study 5).
#

R² R² (variance explained)

Statistics

Also: r2, variance explained

How much of the variation in outcomes a model explains (0 to 1).

Simple
0.76 means the model accounts for about three-quarters of the differences between players' post-move output.
Technical
1 − SSE/SST on the locked test moves.
Why we use it
It complements MAE by showing how much of the spread is captured.
In the research
Final model R² 0.76 vs 0.58 for the naive baseline (Study 5).
#

AUC

Statistics

Also: roc, area under the curve, ranking accuracy

How well a model ranks players who succeed above those who don't (0.5 = coin flip, 1 = perfect).

Simple
Pick one mover who kept his output and one who didn't. AUC is how often the model ranks them the right way round.
Technical
Area under the ROC curve for P(retention ≥ 75%).
Why we use it
It measures ranking quality independent of any cut-off.
In the research
The chance of keeping ≥ 75% ranks players with AUC 0.80 (Study 5).
#

Brier score

Statistics

Also: brier score, probability accuracy

How accurate probability forecasts are (lower is better).

Simple
Saying '90%' and being wrong is punished more than saying '60%' and being wrong.
Technical
Mean squared error of predicted probabilities against 0/1 outcomes.
Why we use it
It checks the probabilities themselves, not just the ranking.
In the research
Brier 0.179 vs 0.224 for a same-for-everyone base rate (Study 5).
#

Modelling

PPML Poisson pseudo-likelihood (PPML)

Modelling

Also: poisson, poisson pseudo-likelihood, regression model

A model for counts and rates that reads results as percentage changes.

Simple
A standard tool for modelling things like 'shots per minute' that handles lots of zeros well.
Technical
Poisson pseudo-maximum-likelihood on per-minute output with minutes as exposure and high-dimensional fixed effects; robust to non-Poisson variance.
Why we use it
It gives effects as clean percentages (e.g. +27%) and stays valid when the data aren't perfectly Poisson.
In the research
The +27% home xG edge is a PPML estimate.
#

Fixed effects Player × team × season fixed effects

Modelling

Also: controls, player team season, within-player

Comparing each player only with himself.

Simple
Instead of comparing Haaland with a full-back, the model only asks: how did this player, at this club, this season, do in one situation versus another?
Technical
A separate intercept for every player × team × season, so identification comes only from within-player variation; errors clustered by match.
Why we use it
It removes differences in ability, team and season, leaving the context effect.
In the research
All headline Study 2 and 3 effects are 'same player, team and season'.
#

Shrinkage Empirical-Bayes shrinkage

Modelling

Also: empirical bayes, reliability adjustment

Pulling noisy individual estimates toward the average by how unreliable they are.

Simple
A player with a tiny sample gets pulled most of the way back to average; a player with lots of data keeps more of his own number.
Technical
Empirical-Bayes posterior means: each player's estimate is weighted against the population mean by its precision relative to the between-player variance (τ²).
Why we use it
It gives the fairest single guess for each player and stops small samples dominating rankings.
In the research
After shrinkage, 0 of 4,665 players have a reliable personal home edge (Study 2).
#

Regression to the mean

Modelling

Also: regression, mean reversion

Extreme results tend to be followed by more ordinary ones.

Simple
The player who looked brilliant at home last season was partly lucky. Next season the luck evens out and he looks more ordinary.
Technical
When a measurement is part signal, part noise, the most extreme observations are disproportionately noise, so they shrink toward the average on re-measurement.
Why we use it
It explains why so many 'special' players disappear when you check them again.
In the research
The 2015–19 'top 10 home players' were exactly average in 2022–25; correlation between periods 0.018 (Study 2).
#

Leakage

Modelling

Also: look-ahead, hindsight, data leakage

When a model accidentally uses information it wouldn't have had at the time.

Simple
Like predicting a match using the final score. It looks brilliant and is useless.
Technical
Any use of post-decision information in features or ratings; tested by scrambling later results and checking nothing earlier changes.
Why we use it
Every prediction here must be one you could actually have made on the day.
In the research
0 changed ratings in 155,282 leakage checks; Understat's post-match 'forecast' was never used.
#

Locked test seasons

Modelling

Also: test set, holdout, unseen seasons, out of sample

Seasons the model never saw until it was finished — scored once.

Simple
Like sealing an exam paper until the student has stopped studying.
Technical
The final model was fixed before 2023/24–2025/26 moves were opened, then scored once on 142 unseen moves.
Why we use it
It's the honest test of whether a model works on the future, not just the past.
In the research
On the locked seasons the final model's error was 0.096 vs 0.129 for 'he'll do what he did' (Study 5).
#

Calibration

Modelling

Also: calibrated, probability check

Whether '70% likely' really happens about 70% of the time.

Simple
A forecaster who says 'rain: 70%' should see rain on about 7 of every 10 such days.
Technical
Observed frequency vs mean predicted probability in bins; the ≥75% threshold is calibrated at the top, lower thresholds were not and are not used.
Why we use it
An uncalibrated probability looks precise but misleads.
In the research
The 50% and 60% retention thresholds were not well calibrated, so the site only shows ≥75%.
#

80% range 80% prediction range

Modelling

Also: prediction interval, prediction range, 80% range, uncertainty

Where the player's output should land 4 times out of 5.

Simple
Not a guarantee: one move in five should land outside it.
Technical
10th–90th percentile of a log-normal predictive distribution with role-group bias and SD from out-of-sample errors.
Why we use it
Individual forecasts are uncertain; the range shows how much.
In the research
A typical 80% range is ±0.2 xG + xA per 90; strikers 4 of 4 and midfielders 5 of 6 landed inside it in the case study.
#

Multiple testing Multiple-testing correction

Modelling

Also: benjamini-hochberg, bh, false discovery, multiple comparisons

Adjusting for the fact that testing many things guarantees some flukes.

Simple
Test 20 weather effects and one will look 'significant' by chance. The correction raises the bar to account for that.
Technical
Benjamini–Hochberg false-discovery-rate correction plus permutation tests.
Why we use it
With thousands of players and metrics, flukes are certain unless corrected.
In the research
0 of 23,022 player × metric opposition estimates survive the correction (Study 3).
#

Permutation test

Modelling

Also: permutation, shuffle test

Shuffling the data many times to see how often chance alone produces a result this big.

Simple
If you shuffle which matches were 'home' and still see the same pattern, the pattern wasn't about home at all.
Technical
Builds the null distribution by randomly permuting labels (e.g. home/away within a player's own matches) and comparing the real statistic against it.
Why we use it
It checks headline effects without leaning on textbook assumptions about the data.
In the research
Benjamini–Hochberg corrections and permutation tests were run for every headline effect (Final Summary, key methods).
#

Placebo test

Modelling

Also: placebo test, fake dates

Re-running an analysis on fake dates to check the effect isn't an artefact.

Simple
If 'no fans' really caused the drop, pretending the fans left in a different year should show nothing.
Technical
E.g. shifting the closed-doors dates 2–4 years earlier; a real crowd effect should vanish on the placebo dates.
Why we use it
It rules out the effect being a quirk of the method or the calendar.
In the research
Moving the closed-doors period 2–4 years earlier found nothing beyond chance (Study 2).
#

Cosine similarity

Modelling

Also: similarity, cosine, distance

How alike two players' profiles are in shape, ignoring overall size.

Simple
Two players who do the same things in the same proportions score high, even if one does more of everything.
Technical
1 − cosine distance between z-scored adjusted profile vectors; checked against Euclidean distance across four profile windows.
Why we use it
It finds players with the same style rather than just the same volume.
In the research
Bruno → Serie A: Samardžić, Dybala and Chukwueze are the robust matches (Study 4).
#

Similarity score Similarity score (closer than X%)

Modelling

Also: similarity, closer than, profile similarity

How close a candidate's profile is to the target's, as 'closer than X% of the pool'.

Simple
A score of 90 means this player's profile is closer to the target than 90% of the players in that role.
Technical
100 × (1 − percentile of the candidate's distance among the query's distances to the role pool).
Why we use it
A raw distance means nothing on its own; ranking it against the pool makes it readable — and it is a resemblance, not a forecast.
In the research
Zirkzee's profile was closer to Højlund's than only 1% of the 2024 striker pool — the least Højlund-like on the board (case study).
#

Role Role (data-derived)

Modelling

Also: position, role pool

A position group defined by what players actually produce, not their listed position.

Simple
A central attacking midfielder turns out to produce like a wide forward, so they're compared in the same pool.
Technical
Ward clustering of Understat's granular positions on per-90 output gives seven roles.
Why we use it
Comparing like with like matters more than the name on the team sheet.
In the research
Bruno's comparison pool includes inside forwards (Study 4).
#

Profile window

Modelling

Also: window, seasons

How much of a player's history the similarity profile uses.

Simple
One season can be a fluke; three seasons is a steadier picture.
Technical
Windows: last three seasons (w3), two seasons (w2), most recent 2,500 minutes, one season (w1). Single seasons are noisy.
Why we use it
Results that hold across several windows are more trustworthy.
In the research
Three-season profiles find the same player in the top 9% of the destination pool after a move (Study 4).
#

Self-retrieval Self-retrieval validation

Modelling

Also: validation, find the same player

Testing similarity by checking it can find the same player again after he moves league.

Simple
If the method is any good, it should recognise a player in his new league.
Technical
Profile before a move is matched against the destination league's role pool after it; the player's own post-move profile should rank near the top.
Why we use it
It validates the method without relying on anyone's opinion of who is 'similar'.
In the research
Three-season profiles find the same player in the top 9% after a league move (Study 4).
#

Recruitment

Retention

Recruitment

Also: survives, kept output

How much of his output a player keeps after a move.

Simple
100% means he produced exactly as much in the new league; 80% means he lost a fifth.
Technical
Output after ÷ output before (xG + xA per 90), also compared with matched stayers to remove regression to the mean.
Why we use it
It's the direct answer to 'will his game travel?'.
In the research
Into the Premier League players keep 82%; out of it, 116% (Study 5).
#

Comparable stayers

Recruitment

Also: comparable stayer, benchmark

Players who didn't move but had the same role, output and age — the fair comparison.

Simple
Players often move after a great season, and great seasons are usually followed by worse ones anyway. Stayers show what would have happened without the move.
Technical
Matched control group at the same club and role with similar pre-period output and age, to net out regression to the mean.
Why we use it
Without them, every move looks costly.
In the research
Movers keep 97% (CI 93–101%) of output relative to comparable stayers (Study 5).
#

Destination league

Recruitment

Also: target league

Player output varies with league context. The similarity model adjusts for league differences before comparing profiles.

Simple
Asking for Serie A means: which Serie A players have profiles that resemble this player once league, opponents and venue are taken out?
Technical
Profiles are adjusted for opponent, venue and league environment, then compared only with players in the same data-derived role in the chosen league.
Why we use it
A raw stat line carries its league with it. Comparing across leagues without adjusting would mostly match players by league, not by style.
In the research
Bruno Fernandes → Serie A: the robust matches are Samardžić, Dybala and Chukwueze (Study 4).
#

League effect League effect (transfer model)

Recruitment

Also: league adjustment, league difficulty

The model's adjustment for how hard the destination (and origin) league is.

Simple
Moving into the Premier League costs a typical player some output; moving into Ligue 1 adds some. The model learns how much from past moves.
Technical
Destination- and origin-league terms in Model A, additive in xG + xA per 90, fitted on earlier moves; relative to the average move.
Why we use it
League difficulty is the biggest single driver of what survives a move. Because the effect is a fixed amount, it over-penalises low-output players — the model's known weakness.
In the research
Into the Premier League: −21% (−30 to −11%) relative to the average move; into Ligue 1: +25% (Study 5).
#

League-difficulty ladder

Recruitment

Also: league difficulty, ladder

One ranking of how hard each league is to produce in.

Simple
Premier League hardest, Serie A and La Liga in the middle, Bundesliga and Ligue 1 easiest. Moving down the ladder boosts numbers; moving up costs them.
Technical
A single league effect per league explains all 20 move directions; pair-specific effects add nothing (p = 0.91).
Why we use it
It means you don't need a separate rule for every pair of leagues.
In the research
Serie A → Premier League keeps ~80%; Premier League → Serie A ~114% (Study 5).
#

Club strength Destination club strength

Recruitment

Also: destination club, club rating, promoted

How strong the team a player joins is.

Simple
Joining a better team usually means more of the ball and better chances.
Technical
The destination club's market-implied rating at the decision date; promoted clubs get an additional term.
Why we use it
It matters as much as the league.
In the research
Weakest third of destination clubs: 79% retention; strongest third: 106% (Study 5).
#

Naive baseline

Recruitment

Also: naive, baseline

The simplest forecast: 'he'll do what he did before'.

Simple
Any model worth using has to beat this.
Technical
Post-move output predicted as equal to pre-move output.
Why we use it
It's the honest benchmark for a transfer model.
In the research
Naive MAE 0.129 vs final model 0.096 on locked seasons (Study 5).
#

Comparables Comparable players

Recruitment

Also: similar players, shortlist

Statistically similar players, used to find candidates.

Simple
A shortlist of players who play like someone you know.
Technical
The top-ranked players by context-adjusted profile similarity in a destination league.
Why we use it
Good for discovery — but they don't forecast how a player will do after a move.
In the research
Comparables alone were worse than the player's own history at predicting post-move output (Study 5).
#

Model A Transfer prediction (Model A)

Recruitment

Also: model a, transfer prediction, prediction, forecast

The Study 5 model that predicts a player's xG + xA per 90 after a league move.

Simple
It starts from what the player has done over several seasons, then adjusts for his age, the club he joins and the league he moves into.
Technical
Regression on player history (three-season level), age, minutes, both clubs' strength, both leagues and a promoted-club term; fitted on earlier moves and fixed before the locked test.
Why we use it
It beats 'he'll do what he did before' on seasons it never saw — and it is the model behind the Transfer Calculator and the Manchester United case study.
In the research
Fixed before the test and scored once: MAE 0.096 on 142 unseen moves, vs 0.129 for the naive baseline (Study 5).
#