Inside the Analysis Room: Why a Football Analyst Must Learn to Say 'Not Enough Data'
core_answer: Phân tích bóng đá chỉ có giá trị khi dữ liệu đầu vào được kiểm chứng qua nguồn, mẫu trận và bối cảnh. Khi thiếu dữ kiện, câu trả lời chuyên nghiệp nhất của nhà phân tích là tuyên bố chưa đủ cơ sở để kết luận, thay vì dựng lên một kết luận nghe có vẻ định lượng.
key_facts: Modrić tại World Cup 2018 được ghi nhận 24 pha nhận bóng giữa các tuyến, phần lớn quãng đường di chuyển không hướng về khung thành đối phương.; Morocco tại World Cup 2022 để Tây Ban Nha chuyền trên 1.000 đường trong một trận nhưng chặn gần hết bóng nguy hiểm vào trung lộ.; Năm 2020, các sân vận động trống không hơn 112 ngày; quy tắc thay người được nới lên 5 suất, gây bất lợi cho các đội pressing tầm cao.; Hè 2024, thương vụ cho mượn Emile Smith Rowe từ Arsenal được thiết kế quanh hệ thống hai tiền vệ lùi.; Ba tầng lọc trước khi công bố gồm kiểm chứng nguồn, kiểm chứng mẫu trận và kiểm chứng hệ quy chiếu.
source_attribution: Phân tích của Kim Jae-sung, blogger chiến thuật tại Liverpool, tổng hợp từ ghi chép theo dõi trận đấu 2018-2024 | Cross-checked: VuaBong.vn
related_qa: q: Vì sao nhà phân tích nên công bố 'chưa đủ dữ liệu' thay vì đưa ra kết luận?, a: Vì khi đầu vào trống, mọi kết luận đều là bịa đặt và phá hủy uy tín chính xác của người phân tích.; q: Chỉ số 'sức bền phòng ngự' đo điều gì?, a: Nó kết hợp quãng đường chạy tốc độ cao với tỷ lệ tắc bóng thành công khi đã mệt, nhằm đo thời gian một cầu thủ còn giữ được cấu trúc.; q: Làm sao kiểm chứng một tin chuyển nhượng trước khi chia sẻ?, a: Cần kiểm tra động cơ của nguồn, kích thước mẫu trận liên quan và độ tương thích chiến thuật với hệ thống của câu lạc bộ.
Inside the Analysis Room: Why a Football Analyst Must Learn to Say "Not Enough Data"
An evening in Liverpool
On my desk in Liverpool sits a thick, worn leather notebook. I started it in the summer of 2026: the positional coordinates of every player, recorded every five minutes, for the matches I watch on screen. Ten years of looking at football through an analytical lens has taught me something simpler than any complex model: a good analyst is not the person who can tell the most stories, but the person who knows exactly when there is not enough basis to tell any story at all.
That Saturday evening, a piece of information spread through a private group chat of analysts I belong to. It was short: a young midfielder, who had not even played fifty matches at the top level, was reported to carry a valuation of one hundred million euros. Within two hours it had appeared on at least seven different platforms. No one could name a source. No one asked where the figure came from, on what sample of matches, by what method.
I spent the whole evening tracing it. The result: it began as a status update from an agent, was aggregated by one account, and then became information at the next steps. No contract, no negotiation, no confirmation from any club. A story built out of nothing, moving faster than the truth exactly the way it does every transfer window.

What made me think was not the rumor itself. It was the speed. From appearance to being treated as a fact, the elapsed time was measured in hours. There was no verification room, no review step, no one standing up to say the number had never existed. And in that gap, an analyst has two choices: feed the story for clicks, or say something nobody wants to hear.
Context: an ecosystem without brakes
Football operates on three parallel streams of information, and they flow at different speeds. The first is on-pitch data — passes, duels, distance covered — something anyone with a patient eye can verify. The second is transfer information, where every party has its own motive to distort: a club wants leverage, an agent wants to inflate value, a seller wants to manufacture scarcity. The third is narrative — how public opinion packages events into a story with a protagonist, an antagonist, and a climax.
These three streams do not move at the same speed. Data flows slowest. Transfers flow faster. Narrative flows fastest, and it usually overtakes both. When a story is told before the data arrives, a reader at the end of the chain has no way to tell a conclusion drawn from observation from a guess dressed up in language that sounds quantitative.
For an analyst, this speed mismatch is the greatest professional trap. It rewards those who speak first and delays the reward for those who wait. In the English market where I work, immediate reward arrives as clicks, shares, and advertising deals. Delayed reward arrives as credibility — something built only by the times you were right, and destroyed by a single time you were wrong.
There is a striking paradox in this reward structure. Credibility takes years to build and an afternoon to destroy. Clicks arrive instantly. When the two are placed side by side, the instant reward always wins in the short term. That is why the market produces so much content that looks like analysis but is really just a story wrapped in numbers.
Anatomy of a fabricated transfer story
Across many transfer windows, I have noticed that fabricated stories follow an almost repeating sequence. Step one, a party with an interest releases vague information, usually a number without conditions. Step two, an aggregating account restates it in more certain language, stripping away the ambiguity. Step three, a larger platform cites the aggregator as if it were the source. Step four, the story becomes a fact simply because it appears in many places.
This mechanism works because readers have no way to trace backward. Once a piece of information has appeared in dozens of places, asking about its origin becomes meaningless to an ordinary reader. They see consensus, and consensus is mistaken for evidence.
What is worrying is that the mechanism exploits the reader's strengths. Readers are busy. They do not have time to cross-check every number. They place their trust in reputation rather than evidence. And reputation, in this case, is measured by how often something appears, not how often it is verified.
For an analyst, this is a direct warning. If I cite a number without checking the source, I am not merely passing on information. I am adding a layer of legitimacy to the story, because readers tend to trust a credentialed person over an anonymous account. Each time I do that, I destroy part of what I am trying to build.
Mechanism: why fake numbers still sound real
There is a curious paradox in how football information is received. A qualitative statement — this player is good — is always challenged for evidence. A quantitative figure, such as a specific progressive-passing metric per ninety minutes, is usually accepted immediately, with no one asking how it was measured.
That asymmetry is an exploitable weakness. Anyone who wants to give a weak claim a professional sheen can wrap it in a number. The number conveys a sense of science without any actual science. That is why, over years of watching the transfer market, my first question is never what the number is, but where it came from.
The three most common sources of noise I encounter:
First, a sample that is too small presented as a rule. Three good matches become consistent form. One nice touch becomes a signature skill. From the habit of recording coordinates every five minutes, I know a sample under ten matches is only enough to form a hypothesis, not to conclude a trend.
Second, a metric detached from context. A defender with a high tackle success rate may simply be playing in a deep-lying defensive side where every intervention is easier. The same metric, placed in two different systems, can carry two opposite meanings.
Third, the motive of the source. An agent has reason to quote a price higher than reality. A buying club has reason to quote a price lower than reality. Checking the source's motive is a step that cannot be skipped, even though it yields no pretty number to quote.
When these three sources of noise combine, the reader receives a product that sounds highly professional but has no foundation. It has tables, jargon, and a confident tone. It lacks only one thing: truth as a basis. Every formation is a hypothesis, and the match is the experiment. But an experiment with no input data, no matter how beautifully presented, yields no conclusion of value.
A lesson from Croatia 2026: every formation is a hypothesis
At eighteen, as a first-year student in Liverpool, I wrote a twelve-part series on the diamond carousel in Croatia's midfield at the 2026 World Cup. I did not write about inspiration. I logged twenty-four receptions by Luka Modrić between the lines, measured the distance he covered, and arrived at a conclusion some considered cold.
Of Modrić's total distance, the portion moving toward the opponent's goal was far smaller than the match felt. He moved a great deal, but most of it was movement to create space for others, to drag defenders out of position, to open a diagonal pass. He was an architect of gaps, not a machine endlessly running forward.

For an analyst, this is the kind of finding that changes how you read a famous player entirely. The brightest star is not the one who runs most, but the one who controls the most space. People see him in every key moment, but do not see that most of his work lies in opening a path for others to enter.
Croatia did not create a miracle; they drew a map. And a map is only correct when read in the right frame of reference. Read Croatia in the frame of who runs most, and you draw a completely different team, one that never existed on the pitch.
My prediction at the time about the erosion of midfield control in extra time, due to accumulated distance across consecutive matches, went in the right direction. It was not a hunch. It was the result of reading a movement model and asking what happens when that model is stretched beyond its limit.
I retell this because it bears directly on the subject at hand. The same match, with the same numbers, can be told as two opposite stories. Honesty does not reside in the number. It resides in whether the analyst is transparent about the frame of reference being used.
Morocco 2026: defending is spatial architecture
Two years later, at the 2026 World Cup, I was hired by a sports channel as a remote analyst. I tracked all six Morocco matches. My job was not to praise spirit. It was to chart.
Morocco's deep 4-3-3 allowed Spain to complete over a thousand passes in a single match, but the number of genuinely dangerous balls into the central lane could be counted on one hand. Morocco's defensive midfield zone occupied the majority of activity time, while Spain was pushed to the flanks. Morocco did not defend with numbers. They turned space into a maze. Every opponent pass led into a narrower corridor, until there was no way out.
There is a detail raw data does not fully tell. Morocco did not defend only with bodies. They defended by pre-deciding where the opponent would go, then closing the door there. That is the difference between pouring concrete and designing a maze: one endures pressure, the other controls it.
Men like Achraf Hakimi and Hakim Ziyech are often cited as symbols of Morocco's spirit. But what interested me more was how the system allocated space to them. Hakimi was free to push high because the flank behind him was covered by defensive midfielders under a clear rule. Ziyech was allowed to hold the ball longer than usual because the runners around him were pre-positioned. No individual shone outside the structure. The structure created the space in which they shone.
Before the semi-final against France, I made a prediction unwelcome to Morocco's admirers. Based on accumulated defensive actions — total high-speed running accumulated across six matches, the highest in the tournament — I argued their capacity to absorb pressure had reached its threshold. The result matched the script: defeat, and defeat not from a sudden loss of form but from the cumulative cost of a defensive system demanding relentless motion.
From this I built a personal metric, defensive endurance: high-speed distance combined with tackle success rate while fatigued. This metric does not measure how strong a player is. It measures how long a player can maintain structure. For a deep-defending side, that is the right question.
112 days and the disappearance of noise
In 2026, stadiums stood empty for more than one hundred and twelve days. As someone attached to process, I decided to dissect the effect of losing the wall of noise. I analyzed Liverpool's home matches without spectators and found something that made me recheck several times: their high defensive line committed far more positional errors than with a crowd present.
The cause was not skill. It was signal. With a crowd, the roar works as a community warning system: it compresses the reaction time of the whole shape. When the noise disappears, midfielders lose a channel of information, and the space behind a high line becomes more dangerous.
At the same time, substitutions were expanded to five. This change has far greater impact than the way it is usually discussed. High-pressing teams suffer clearly: when opponents can make five changes, fresh energy injected in the second half directly undermines their pressing model. That is why I wrote that one hundred and twelve days without football was not merely an emotional void. It was a natural test of the hypothesis about the crowd's role in tactics.
112 days without football, and the substitution rule is a lifeline. But that lifeline saves one team and sinks another, depending on their philosophy. A rule change is not fair or unfair to all alike. It exposes the fit between the rule and the playing philosophy.
From that experience, I added to my pre-match checklist factors beyond the pitch: crowd, substitution regime, travel schedule, rest days between matches. To an analyst reading only on-pitch data, these are invisible. To an analyst reading context, they are part of the match.
Smith Rowe 2026: checking fit before publishing
In the summer of 2026, while working for a football media startup in Liverpool, I covered the transfer window. Through a scout, I was among the first to know about Emile Smith Rowe's loan move from Arsenal — a deal designed around a double-pivot system.
What I checked before publishing was not the player's value. It was the fit. Smith Rowe receives a substantial volume of passes in the left half-space, exactly the zone the new system needed to exploit. On the numbers, it was a snug fit. On process, it was a lesson: before saying a club bought a good player, ask what problem their system needs solved.
The transfer market does not buy players; it buys problems. An expensive signing is not wrong because of its price. It is wrong when it solves a problem the club does not have. And a cheap signing is not right because of its price. It is right when it fits the system, even if nobody notices at the time of signing.
This is where I disagree with how the market prices young talent. The youth price bubble is inflating, and it inflates on potential rather than evidence. One hundred million euros for a player who has not played fifty top-flight matches is a naked gamble, even when dressed in the language of strategic investment. The issue is not paying a lot. The issue is that the price is not anchored to any verifiable fact.
From that experience, I added a rule to my daily work: before publishing any transfer information, check the source's motive and the tactical fit. These two checks do not guarantee I am right. They only remove the times I am certainly wrong without knowing it.
Counter-metrics: what I force myself to hunt for
There is a habit I consider the boundary between analysis and propaganda: for every claim, I look for at least one metric capable of contradicting it before writing.
If I intend to say a defensive midfielder is playing well, I look for how many times he was dribbled past. If I intend to say a pressing team is effective, I look for the number of long balls they are forced to absorb after a failed press. Selective data can always tell a pretty story. The analyst's job is to find the number that makes that story harder to keep pretty.
This is the hardest part of the work, because it runs against natural instinct. When you have spent hours building an argument, hunting for evidence against it feels like dismantling a machine you just assembled. But that step is what separates someone who presents statistics from someone who analyzes.
Before praising the star, measure the gap he leaves behind. A brilliant playmaker may be masking the fact that his midfield is constantly pierced. A centre-back scoring from corners may be trading that for lost positions when his team counter-attacks. If you measure only the praised part, you are reading half a match and calling it the whole.
I do not believe in randomness; I believe in repeated passes. A single beautiful moment can be luck. A pattern repeated across many matches is a tactical fact. My job is to distinguish the two, and that can only be done if I force myself to seek evidence that does not support my conclusion.
The blind spot: when systems cannot explain people
Every model has an edge. I have to remind myself of this often, because I am the kind of person who believes in structure.
Metrics answer what happened, not why. If a player misses in the eighty-eighth minute, the data tells me position, angle, shot power, and the goal probability of that shot. It does not tell me how much he slept the night before, what is on his mind, or what week he has had. Fitness and mental state are not in the model, and cannot be replaced by a single variable.
I have written that demanding a player prove himself in his first match back from injury is cruel, and that it raises the risk of re-injury. This is not a sentimental judgment. It is an application of the same principle: the model is based on a body's normal state, while a body after injury is not in a normal state. Applying a normal model to an abnormal body is using the wrong tool.
So when I analyze, I always state the edge of the model I am using. Metrics can answer what happened. They cannot answer how a human being felt on a particular night. Without acknowledging that edge, analysis becomes fabricated certainty.
I also have to remind myself that a fondness for system metaphors — maps, mazes, architecture — can pull a piece too far from the pitch. Whenever I write a passage about structure, I attach it immediately to a concrete situation: a passage of play, a player, a substitution decision. If I cannot do that, I am writing about a match that does not exist.
There is another temptation to watch for: drowning in detail and losing the narrative thread. Accurate data with no thread leaves the reader behind. So after each main point, I write a summary sentence in plain language, an anchor for the reader before moving on.
VAR and the editing of matches
On refereeing, I hold a position formed over years of observation: drawing offside lines to the millimetre is eroding football's attacking instinct.
The problem is not the technology. It is how the technology is used to redefine what counts as correct. A striker who runs beyond a defender by half a footstep — a movement the naked eye cannot detect — is now ruled offside because of a digital line. The instinct to time a run, honed since childhood, becomes meaningless against software.
This changes player behaviour at a deeper level than individual decisions. When a player knows that running half a step early can be detected, they tend to run later for safety. Running later means fewer dangerous balls. Fewer dangerous balls means fewer goals. Here, technology does not merely correct an error. It reshapes how the game is played.
The referee becomes an editor of the match, not just an administrator of the laws. And that is the least-discussed consequence in VAR debates. We argue about the accuracy of individual decisions, rarely asking about the cumulative effect of thousands of accurate decisions on the nature of the sport.
Three filters before publishing
Over the years, I have reduced my work to three filters, applied to every piece.
The first is source verification. Who said this, what do they gain if I believe them, and how often have they been right before? For transfer news, this question eliminates most circulating content.
The second is sample verification. How many matches does this claim rest on, under what circumstances, against representative opponents? A sample of three matches against three weak opponents is not enough to describe a trend.
The third is frame-of-reference verification. By what standard am I measuring this player, and does that standard fit his role in the system?
These three filters are slow. They make me publish less than others. But they stop me from becoming a link in the chain of fabricated transmission — a chain that grows a little longer every transfer window.
The counter-intuitive angle: sometimes the right answer is to say nothing
There is a professional paradox I took years to accept. In an industry that rewards speaking, silence is treated as failure. No article, no views, no contract. But there are moments when the most professional output is a blank page.
When the input contains nothing — no information points, no named entity, no source, no viewpoint — then every conclusion is fabrication. A genuine analyst must be able to look at an empty input and say: there is not enough basis to assess.
This runs against the instinct of both writer and reader. The reader wants an answer. The writer wants recognition. That pressure produces a dangerous kind of product: content that sounds like analysis but is built on a void. It uses the right jargon, the right table structures, the right confident tone. It lacks only one thing: truth as a basis.
I have refused to write many times for this reason. Not because I was not good enough, but because there was not enough data. It is the kind of decision that generates no clicks, wins no praise, gets no shares. But it protects the only thing an analyst truly owns: a reputation for accuracy.
Tactics are the only thing that cannot be faked on the pitch. Everything else — transfer fees, insider news, dressing-room stories — can be embellished. But when the ball rolls, a shape either fits the idea or it does not. This is why I always return to the pitch as the final anchor. There, every lie must pay for itself in points.
Takeaway: a question for the next match
When you read a football claim in the coming weeks, try asking one question before believing it: if you strip away all the numbers, does the claim still stand? If the answer is no, perhaps the number was there only to convey certainty, not to prove anything.
And when an analyst tells you they do not have enough data, consider reading that as a sign of honesty. In an industry that rewards speed, the person willing to slow down is often the one holding something more valuable than clicks: the ability to distinguish between what they know and what they merely want to believe.
The next transfer window will bring hundreds of new numbers. The question is not which number is right. The question is how many seconds you will spend checking before you share it.
