Football/ Ayush Pawar

Studies 1–3 asked: can we remove context?Study 4 asks: now that we've adjusted for context, can we compare players?

04 / Similarity Complete

Who actually plays like whom?

Can we find functionally similar players across leagues?

  • 9Profile metrics
  • 7Roles from the data
  • 5Leagues

Finding

Context-adjusted profiles find sensible matches across leagues — Bruno → Serie A: Samardžić, Dybala, Chukwueze — but nobody matches his chance creation. Similarity describes a player; it is not a forecast.

02Why it matters

Recruitment often starts with “who plays like him, in a league we can afford?”. Raw numbers can't answer it across leagues: the league, the opponents and the venue all leave fingerprints on a stat line.

Studies 1–3 measured those fingerprints. Now they can be removed.

Once context is removed, can we find a player's functional twin in another league?

03What we did

Chance creation

Involvement

Shooting

Profile

  • Headed share
  • In-box share
  1. Nine-metric profile
  2. Adjusted for opponent, venue and league
  3. Normalized within role
  4. Cosine similarity

Each profile is adjusted for opponent, venue and league, normalized within , then compared using .

Technical detailHow the estimate is built

Roles come from the data: clustering Understat's granular positions on per-90 output gives seven roles. A central attacking midfielder produces like a wide forward, not like other central midfielders.

Metrics are within role and compared by cosine similarity; three-season are the steadiest.

Validated without opinions by : after a league move, does a player's new profile find his old one?

04What we found

The profile finds the same player again — even after a move.

Three-season profiles find the same player in the top 9% of a destination league's role pool after a league move. Adjusted inputs beat raw ones for players who changed league.

After a league move, how high does a player rank against himself?

Median percentile of the player's own earlier profile in the destination role pool. Lower is better.

With three-season adjusted profiles, the player's own earlier profile sits in the top 9% of the destination pool; raw inputs do worse at every window.

Study 04 · self-retrieval on 273 three-season moves

Single seasons are noisy profiles; longer windows are steadier.

We tried to find Bruno's replacement.

Bruno Fernandes's chance creation is 3.6 SD above his role average. Serie A's best comparable is 1.9 SD.

Chance creation above role average

Bruno against the four best creators in Serie A's attacking-mid and wide-forward pool. Standard deviations.

Bruno is 3.6 SD above his role average; the best in Serie A, Martin Baturina, is 1.9. His closest profile matches create less still.

Study 04 · Bruno case study, profiles to 2025/26

Closest profiles in Serie A

  1. 01Lazar SamardžićAtalanta · robustClosest on shooting volume, shot profile; lower chance creation.99
  2. 02Paulo DybalaRoma · robustClosest on involvement; lower chance creation, higher shooting volume, higher shot profile.97
  3. 03Samuel ChukwuezeAC Milan · robustClosest on involvement, shooting volume; lower chance creation, higher shot profile.94

None was a true like-for-like Bruno equivalent.

Samardžić, Dybala and Chukwueze are robust matches on the rest of the profile — but every one of them creates less. A high score means two profiles look alike, not that one player is the other.

About half of what makes a player distinctive survives a move.

After a real move, the player's own pre-move profile forecasts his new-league output as well as his comparables do.

05What surprised us

  1. What the raw data suggested

    A central attacking midfielder sounds like a kind of central midfielder.

  2. What happened after we checked

    On the numbers he produces like a wide forward — so roles were rebuilt from what players produce, not where the team sheet puts them.

  3. Lesson

    A position label describes where a player stands. Output describes what he does.

06What it means

  • Similarity is a discovery tool: it finds candidates worth watching in leagues you don't follow.
  • It is not a forecasting tool. A high score says two profiles look alike — not that one player will produce like the other after a move.

07Limitations

  • Output only: the profile cannot see defending, pressing, pace or physical data.
  • Single seasons are noisy profiles; longer windows are steadier.
  • Study 4's user-curated “known similar pairs” check is still open.
  • Similarity scores rank players within a role pool; they are not probabilities.

SourcesUnderstat · football-data.co.uk · Open-Meteo ERA5 · Wikipedia/Wikidata · OpenStreetMap
Analysis periodProfiles to 2025/26 · frozen dataset · Methodology

08The next question

We had found players who looked similar.

But does similarity actually tell us what happens next?

05 / Transferability