Between the Data Table and the Breath of the Match: The Craft of Reading Esports Data
Q: Phân tích dữ liệu esports là gì và vì sao nó quan trọng? A: Phân tích dữ liệu esports là việc thu thập, xử lý và diễn giải các chỉ số sinh ra từ trận đấu để đánh giá sức mạnh, chiến thuật và tiềm năng của đội tuyển và người chơi. Key facts: - Esports đã chuyển từ ba chỉ số thô (mạng, vàng, trụ) sang hàng nghìn điểm dữ liệu mỗi phút, gồm chênh lệch vàng phút 15, tỷ lệ tham gia hạ gục và điểm tầm nhìn. - Mỗi bản cập nhật patch tạo độ trễ nhận thức kéo dài khoảng hai tuần, khiến dữ liệu cũ mất giá trị dự đoán trong giai đoạn điều chỉnh. - Định giá chuyển nhượng người chơi esports không có thị trường tập trung, không hợp đồng chuẩn và không cơ quan định giá độc lập. - Phân tích dữ liệu cũng được dùng để giám sát liêm chính thi đấu, nhưng tồn tại cả dương tính giả và âm tính giả. - Yếu tố tinh thần người chơi không được bất kỳ chỉ số nào đo lường đầy đủ. Source: Phân tích của Hoàng Việt, Nhà báo dữ liệu tại Shenzhen, đăng ngày 13 tháng 8 năm 2026. Related Q&A: Q: Vì sao chênh lệch vàng sớm không quyết định kết quả trận đấu? A: Vì cùng một chênh lệch vàng có thể phản ánh hai loại lợi thế khác nhau về chất lượng, và tương quan không đồng nghĩa với quan hệ nhân quả. Q: Nhà phân tích dữ liệu esports nên làm gì khi mẫu dữ liệu nhỏ? A: Cần kiểm tra cỡ mẫu, phân tách theo patch và thừa nhận rõ ràng mức độ bất định thay vì gộp dữ liệu qua nhiều phiên bản meta khác nhau.
In an analysis room in Shanghai, the screen displays the gold difference at minute 15: the home team leads by 2,400 gold. According to my model, that number corresponds to roughly a 78% win probability. When the match ends, the leading team has lost. I stay behind, rewind the footage twelve times, and realize what the scoreboard never showed: two missed skirmishes in mid lane, three vision wards cleared, and one badly timed objective call. The data said that team should have won. The match said that team lost. Both are true, and both are incomplete. That is the first lesson I learned entering the trade of reading esports data, and it is a lesson I still relearn every week.
I am a data journalist born in Vietnam and now working in China. My job is to cover esports for an audience already deeply familiar with numbers — a market where every match is analyzed down to the second, every player is priced by an advanced stat sheet, and every transfer window is a quantified marketplace measured to the last yuan. I belong to the kind that colleagues call the Data Monk — the one who tells stories with data. Every piece I write begins with a number, but it never ends there.
The major season is at its peak. Audiences follow the flags and the narratives. But behind every highlight on screen, another layer is unfolding: the battle to name the numbers. Whoever controls the definition of a metric controls the story. And the story of esports data in 2026 is far more complicated than any weekly power ranking suggests.
This piece is a journey inside the craft of reading esports data — not to celebrate the power of algorithms, but to understand why numbers never lie, yet never tell the whole truth either.
Context: When esports learned to count
To understand why esports data analysis became a billion-dollar industry, one must look back at the road already traveled. Ten years ago, when I first started blogging about football and trying to calculate xG from shot data, esports was still in its statistical infancy. People counted kills, counted gold, counted towers. Those three numbers generated most commentary. Today, a professional League of Legends match can produce thousands of data points per minute, from player positions second by second to the economic value of each vision play.
That explosion did not come from nowhere. It came from three converging streams. First, the maturation of the esports industry, as clubs shifted from ad hoc team models to professional organizations with coaching staff, analyst departments, and data science units. Second, the spread of official APIs allowing match data collection with unprecedented granularity. Third, and most important to me, a shift in thinking: from narrating the match to narrating through the context of the match.
When a metric is born, it is not merely a measurement. It is a stance. The gold difference at minute 15 was created by people who believe the laning phase decides the match's fate. Kill participation was created by people who believe collective contribution outweighs individual stats. Vision score was created by people who believe information warfare is the root of every victory. Every stat sheet carries within it an ideology about how electronic football should be played. And the battle to name the numbers is precisely the battle between those ideologies.
I witnessed this shift early. In 2026, when the World Cup kicked off in Russia, I had just turned 18 and was a first-year student in Shenzhen. In the semifinal between France and Belgium, I calculated France's xG at roughly 1.6 and Belgium's at 0.8, but France won 1-0 through Umtiti's header from a set piece. My model said France did not deserve to win by that much. The match said France won. I spent a full month reviewing all the footage, analyzing each play, and adjusting the model to add weight for set-piece situations. The subsequent article was more accurate. But the lesson was not about accuracy. The lesson was this: data has limits, and the writer who uses data must be the first to admit them.
When I moved from football to esports, I carried that principle with me. And I discovered that esports is the ideal environment to test it, because esports generates data naturally. Football has 90 minutes with hundreds of discrete events. League of Legends has 30 minutes with thousands of continuous events. If in football data is a window into the match, in esports data is almost the match itself digitized. That is an opportunity, and also a trap.
Core analysis: The map of metrics and its blind spots
When an esports team walks into the analysis room after a match, they do not look at the scoreboard. They look at a layered map of metrics. The lowest layer is raw stats: kills, deaths, assists, gold, minions. The middle layer is derived stats: gold difference at 15, kill participation rate, damage per gold, vision score per minute. The highest layer is composite stats generated by models: real-time win probability, marginal contribution, advantage conversion efficiency.
Each layer has its own value, and each layer has its own blind spots. Raw stats are easy to understand but misleading. A bot laner with many kills is not necessarily playing better than a top laner with few kills but control of tempo. Derived stats have more context but depend on definitions. Composite stats are the most powerful but also the most opaque — readers do not know what the model is weighing.
I often explain this to readers with an example. Suppose two teams have the same gold difference at minute 15: 1,500. Team A achieved it through three top-lane kills. Team B achieved it through minion advantage and stable vision across all three lanes. On paper, the two are equal. In reality, Team A depends on a hot spot that may cool, while Team B accumulates a structural advantage. The same gold lead, but a different quality of advantage — and the quality of advantage lives in no cell of any spreadsheet.
This is why I never make absolute claims based on a single metric. Whenever I write an analysis, I follow an implicit rule: for each key number, there must be at least two cross-checking sources, and at least one paragraph explaining why that number might be wrong. My readers do not need conclusions so certain they cannot be challenged. They need conclusions so honest they can be verified.
Meta and the patch problem
In esports, data analysis cannot be separated from patch analysis. Each update is a restructuring of priorities. A buffed champion can push pick rate from 5% to 40% in a single week. An adjusted item can erase an entire playstyle. To the analyst, the patch is the strongest independent variable, stronger than form or roster.
But the patch has a property that data struggles to keep up with: cognitive lag. When an update is released, teams need time to explore the new meta. During that window, data from old matches no longer represents true strength. A model trained on pre-patch data can predict wrong across the board — not because the model is poor, but because the world has changed.

From my experience following professional matches, I always note the patch release date next to the match date. If the gap is under two weeks, I mark the data as "in the adjustment period" and downweight it in my conclusions. That is not evasion. It is respect for the reality that the meta is not a constant but a process.
There is a paradox in patch analysis I want to emphasize: a patch does not only change how the game is played, it changes how the playing is measured. When match pace rises, gold at 15 becomes less important and early objective control becomes more important. Analysts must constantly rewrite their own toolkits. Whoever forgets that will forever analyze last season's match.
Regional tactics: Four styles, one meta
One of the most fascinating parts of this job is observing how different regions interpret the same meta. The Korean league is famous for macro play, tempo control, and vision discipline. The Chinese league is famous for fight intensity, top-side pressure, and the ability to create sudden breaks. Europe brings creativity in draft and unconventional compositions. North America, though often underrated, has a talent development system worth studying.
When I compare data across regions, I am always careful of one trap: comparing absolute metrics across different competitive environments. A player with high CS per minute in region A is not necessarily better than one with lower CS in region B, because match pace, opponent quality, and team philosophy differ. Cross-region comparison only makes sense when normalized for context. And that normalization is never perfect.
Numbers do not lie, they simply never tell the whole truth. A high metric can signal talent, or signal a system that favors that individual. A low metric can signal weakness, or signal a sacrificial role. The good analyst distinguishes the two possibilities before drawing conclusions.
Player valuation: When a number becomes a life
No field sees the battle to name numbers as fiercely as the transfer market. When a club signs a player for a reportedly record fee, that number instantly becomes the yardstick of expectation. Fans compare it to others. Media turns it into headlines. The coaching staff has to live with it every day.
But esports player valuation is more complicated than football valuation, because there is no centralized transfer market, no standard contract, and no independent valuation body. The final number is usually the result of a negotiation the public never sees. Every transfer number is a life converted into value. Behind it are age, potential, performance pressure, and non-sporting factors like commercial image and media appeal.
I once wrote an analysis of a controversial transfer, trying to separate the player's competitive value from his commercial value. Reader response split into two camps. One argued I diminished the player by converting a career into money. The other argued I did not go far enough in exposing the market's irrationality. I kept the piece unchanged, but added a methodological note explaining that every valuation is a conditional assumption, and conditions can change.
The lesson: when writing about money in esports, write about people first. Numbers are tools to tell the story, not substitutes for it.
Tournament structure: Where format shapes statistics
An underdiscussed aspect of esports data analysis is the impact of tournament format on the data itself. A double-elimination bracket generates more matches for strong teams, making the sample richer. A single round-robin generates a smaller sample, making metrics less stable. The Swiss system balances opponents but produces fewer direct matchups between top teams.
When I analyze a tournament, I always check sample size before trusting any trend. A team winning five straight group-stage matches could signal strength, or signal an easy schedule. Without format context, a number is an isolated piece. Whether a stadium has fans or not, the match still needs someone to tell it. And that teller must know whether they are telling a group-stage, playoff, or final story.
I also notice an interesting phenomenon: teams show clearly different metrics between group stage and playoffs. Playoffs compress pressure, making teams play safer, commit fewer errors, and often drag matches longer. If you analyze a team only from group-stage data, you can mispredict their playoff behavior. This is one area where pure data models often fail before the eyes of an experienced coach.
Club economics: Cash flow and growth limits
To understand why esports data matters more and more, look at club economics. A professional team has several revenue streams: sponsorship, publisher revenue sharing, prize money, in-game item sales, and streaming revenue. But costs rise too: player salaries, coaching salaries, facility costs, travel costs.
In that context, data analysis becomes a risk-management tool. A club cannot spend infinitely, so it needs to know which dollar yields the highest competitive value. That is why analytics departments have been heavily invested in recent years. They help coaches prepare for opponents, but they also help leadership decide whom to buy, sell, and keep.
However, I always warn about the trap of over-optimizing data in governance. When every decision rests on metrics, clubs can ignore the unmeasurable: team chemistry, fighting spirit, leadership in the room, and fan connection. Those factors appear in no spreadsheet, yet they determine long-term success no less than numbers.
Data is a monastery, but I choose to leave the gate to find the game. I wrote that for football, but it holds for esports too. The analysis room is a safe shelter from the chaos of the match. But the match is where truth lives.
Governance and competitive integrity: The gray zone of numbers
One area where esports data is increasingly used is competitive integrity monitoring. Algorithms can detect anomalous behavior patterns: skewed betting odds, in-game behavior inconsistent with normal form, or suspicious cooperation patterns between teams. This is one of the most ethically charged applications of data analysis.
But here too, the line between analysis and accusation becomes fragile. A fraud-detection model will have both false positives and false negatives. A player having a bad match can be wrongly suspected. A sophisticated cheater can pass every filter. The data journalist must be extremely careful reporting on these topics, because a false accusation can destroy a person's career.
I once refused to write a piece about a player suspected of match-fixing, even though I had data showing anomalous patterns. I refused not out of fear, but because I could not distinguish between signs of fraud and signs of a player in a form crisis. Not enough evidence to conclude. And when there is not enough evidence, silence is a responsible choice.
That is a principle I learned very early: responsible intervention does not mean saying everything you know. It means saying what you can prove, and admitting what you cannot.
A counterintuitive angle: Correlation is not causation
Now comes my favorite part of every analysis: going against popular belief. In esports, a belief is deeply rooted in the community that advanced metrics can predict outcomes with near-perfect accuracy. Power rankings, win-probability models, contribution stats — all create a sense of certainty that we have grasped the rules of the game.
But I want to pose a hard question: do those metrics really measure what we think, or do they merely reproduce what we already believe?
Consider the relationship between early gold difference and win rate. There is certainly a correlation. But that correlation can be explained two ways. Way one: the team with more gold has more resources, so it wins. Way two: the stronger team tends to generate more gold and also wins more, so gold is only a marker of overall strength, not the cause of victory. The two explanations lead to different conclusions about optimal play. If way one is right, teams should focus on accumulating gold. If way two is right, teams should focus on raising overall strength and let gold come.
Most analyses I read implicitly choose way one without verification. That is a major blind spot. 0.35 is a number, but the battle to name it is the truth. The same number can be assigned different meanings, and the chosen meaning often reflects the interests of the chooser.
The paradox of large samples
Another counterintuitive point: a large sample is not always better than a small one. In esports, the number of matches in each new meta is often very limited. If we pool data across many patches, we get a large sample but a sample that blends many different worlds. If we split data by patch, we get a clean but small sample. This is a trade-off with no perfect solution.
Many prediction models I see on the market choose pooling, because it yields prettier, more stable numbers. But a pretty number can be a sign of overfitting to the past, not predictive power for the future. A model trained on three years of data can predict past matches accurately yet fail entirely on the next one, because the meta changed after the latest update.
The biggest blind spot: People
And here is the biggest blind spot of all esports data analysis: the mental state of players. No metric measures the moment a player loses confidence after three straight deaths. No algorithm captures fatigue in a forty-minute fifth game. No dataset records the pressure on a player who knows this may be his last match in national colors.
I once followed a Southeast Asian national team through a regional tournament. Pre-tournament data showed strong attack metrics, stable objective control, and a formidable bot duo. But when the tournament began, they played below their true strength. I sat in the arena, watching players' expressions after each loss: silence, eyes down at the keyboard, shoulders slumped. What happened on screen was not a tactical problem. It was a problem of belief.
After the tournament, I wrote an analysis, but I decided not to center it on the stat sheet. I began with the atmosphere of the arena, with the slowing clatter of keyboards, with the silence between objective calls. Readers responded that this piece made them understand that team more than any previous one. Not because it was smarter, but because it was more honest.
What data cannot measure this moment? I always end every piece with that question. If I cannot answer it, I cut the number that props up the story, because it does not deserve to stand at the center.
A progressive conclusion: Signals of the next round
What is unfolding in esports data analysis is not a race to perfect accuracy. It is a shift from measurement to understanding. Future teams will not only have analytics departments. They will have data storytellers — people who know how to turn a stat sheet into tactical meaning, and know when to put the stat sheet away to listen to the match.
For journalists like me, the challenge of the coming season is to do the opposite of instinct. Instead of adding more numbers, add more context to the numbers we already have. Instead of asserting more firmly, admit uncertainty more clearly. Instead of optimizing for spread, optimize for trust.

I have stood in the empty arena and heard the background hum of esports. It is the sound of thousands of data streams flowing through a match, the sound of decisions made in an instant, the sound of people carrying a number the public calls expectation. My task is not to make that hum louder. My task is to help readers hear clearly what is being said, and to recognize what is being ignored.
The next round will begin with a new update. The meta will shift. Familiar numbers will take on new meanings. And the biggest question I carry into the new season is not which team will win it all, but how we will name their victory. Between a perfect stat sheet and an imperfect match, I know where I will look. Not because the stat sheet is wrong. But because the match is the only place where truth is played, not calculated.
