US Open 2026: where the models were right, and where they missed

By Kate Richardson • September 18, 2026

tennis-ball-and-racket-with-night-city-skyline-in-background

Tennis forecasting has become much better at describing uncertainty, but it has not become a machine for eliminating it. That distinction matters heading into the 2026 US Open, scheduled from August 23 to September 13 in New York. By then, fans will already have seen thousands of percentages, projected draws, Elo ratings and machine-learning forecasts attached to the field.

Those numbers are useful. They are also easy to misunderstand.

A model does not need to name the eventual champion to have performed well, and a correct pick does not necessarily prove that a model was good. The better question is whether the probabilities were calibrated. Did the model react quickly enough to injuries, surface changes and recent form? Did it distinguish between a strong player and a strong player in a difficult draw?

The 2026 season has already provided unusually clear examples. Some major results have rewarded the traditional inputs prediction systems trust most. Others have exposed blind spots that remain even when the data set is deep and the mathematics sophisticated.

The easy part: identifying the strongest players

At the top of the men’s game, models have had several moments that look almost reassuringly logical.

Carlos Alcaraz entered the Australian Open as the No. 1 seed and left Melbourne with the title, beating Novak Djokovic in the final and completing the career Grand Slam. That is exactly the kind of outcome rating systems are built to recognize: an elite player, proven across surfaces, converting a strong pre-tournament position into a major title.

Alexander Zverev’s Roland-Garros victory offered another example. He was the No. 2 seed in Paris and, once several dangerous names disappeared, his path became increasingly attractive from a probabilistic point of view. He still had to survive a five-set final against Flavio Cobolli, but the broad model logic held.

Then came Wimbledon. Jannik Sinner, the top seed, defended his title by beating Zverev in four sets in the final. Again, the predictive ingredients were familiar: elite baseline quality, reliable serving, recent Grand Slam success and a rating profile that had already separated him from most of the field.

If a forecasting system consistently placed Alcaraz, Zverev and Sinner among the highest-probability contenders, it was doing something right. The lesson is not that favorites always win. It is that sustained strength still matters enormously in best-of-five tennis, where the longer format gives the superior player more time to recover from a slow set or a brief tactical problem.

Where the percentages become more interesting

The women’s majors have been a sharper reminder that probability is not destiny.

Elena Rybakina won the Australian Open after coming back from 3-0 down in the deciding set of the final against world No. 1 Aryna Sabalenka. Rybakina was hardly an unknown outsider; her serve, first-strike tennis and previous major pedigree made her a plausible champion in any serious model. But a forecast that simply treated the top-ranked player as overwhelmingly superior would have been too crude.

Rankings are descriptive. Prediction models are supposed to be conditional.

A player may be ranked higher because of results accumulated over many months, while a match model can place greater weight on hard-court performance, serve effectiveness, return numbers, opponent quality and recent form. The best systems do not ask only, “Who has been better this year?” They ask, “Who is more likely to be better in this particular match, under these conditions?”

The difference became even more obvious at Roland-Garros.

Mirra Andreeva, the No. 8 seed, won her first major at 19. Her victory itself was not an absurd result for a data-driven model: she had already reached the semifinal in Paris earlier in her career. The extraordinary part was the identity of the finalist. Maja Chwalinska, a qualifier ranked No. 114, came through qualifying and reached the championship match.

No pre-tournament model sensibly built around ranking, recent tour-level form and opponent-adjusted strength would have assigned Chwalinska anything close to a leading title probability. Nor should it have. Her run is not evidence that probability models are useless. It is evidence that low-probability events sometimes happen.

That sounds obvious, but it is the point most often lost when predictions are judged after the result is known.

A correct forecast can still be a bad forecast

Suppose a model gives a favorite a 90 percent chance of winning and the player survives only after saving match points. The prediction was technically correct, but the confidence may have been excessive. Another model might give the same player 64 percent. Both selected the winner, yet the second could still have been better calibrated.

This is why accuracy alone is a limited way to judge tennis predictions.

A large Journal of Sports Analytics study examined more than 25,000 ATP matches and almost 14,000 WTA matches from 2010 through 2019. Its broader conclusion was sobering for anyone selling certainty: across a wide range of machine-learning approaches, average prediction accuracy was difficult to push much beyond roughly 70 percent, while betting-market information already contained much of the useful signal.

Tennis still contains enough variance, matchup effects and short-term physical uncertainty that a meaningful share of outcomes will remain genuinely difficult. The practical standard is not perfection, but calibration, responsiveness and the quality of the variables.

Why serve and return data travel well to New York

Hard-court tennis gives forecasters a relatively clean environment for combining long-term ratings with point-level performance.

Serve points won, return points won, first-serve effectiveness and opponent-adjusted performance can reveal changes that rankings need longer to absorb. Research on forecasting serve performance has also shown the value of correcting for sample size and strength of schedule rather than treating raw percentages as if every opponent were equally difficult.

A player may arrive in New York with a 9-2 hard-court record built largely against opponents outside the top 50. Another may be 6-4 after facing four top-10 players and several elite returners. A simple recent-win percentage favors the first player. An opponent-adjusted system may reasonably prefer the second.

This is where services such as TennisPredictions.ai fit naturally into the modern tennis conversation: not as substitutes for uncertainty, but as tools that can combine several pieces of evidence into one probability. The useful number is not the one that looks most certain. It is the one that best reflects what is known before the match begins.

That also explains why model outputs should move when information changes. A late withdrawal, a physical problem visible in the previous round, a difficult five-set recovery or a sudden drop in second-serve performance should matter more than a static rating generated days earlier.

Wimbledon showed the blind spot in plain sight

If the men’s tournament at Wimbledon rewarded the favorite structure, the women’s event produced one of the clearest warnings of 2026.

Linda Noskova won the title at 21, defeating Karolina Muchova in an all-Czech final. She had already won Berlin before Wimbledon and, by the end of the fortnight, had collected more tour-level grass wins in 2026 than any other woman. In retrospect, the signs look obvious.

They were not obvious enough beforehand to make her the consensus favorite.

This is a classic forecasting problem. Models are often strongest at measuring established level and weaker at identifying when a player is moving rapidly from one performance tier to another. A young player with improving serve numbers, better movement and growing confidence may still carry a rating partly shaped by matches played before that improvement.

The faster the change, the greater the lag.

Noskova’s run also demonstrated the limits of treating tournament forecasts as seven independent match predictions. She saved a match point against Sorana Cirstea in the third round. If that point goes the other way, every later projection becomes irrelevant. Even a player favored in every individual round can have a modest chance of winning seven consecutive matches because the probabilities multiply.

That is why an outright probability of 18 or 22 percent can make someone a clear title favorite at a Grand Slam while still meaning that player is more likely to lose the tournament than win it.

Draw structure is not a footnote

One of the most persistent errors in public tennis forecasting is to discuss title chances as if the draw were neutral.

Two players with almost identical underlying strength can have very different championship probabilities if one is placed near a dangerous unseeded opponent, a former champion returning from injury or a high-variance server capable of shortening sets into tiebreaks.

A player rating is therefore not yet a tournament model.

To estimate US Open title chances properly, a system has to combine match probabilities with the actual bracket and account for the many possible paths through it. This is often handled by simulation: the tournament is played virtually thousands of times, with each match sampled according to its estimated win probability. The share of simulated tournaments won by each player becomes the title probability.

The method is intuitive, but the answer is only as good as the probabilities underneath it. Simulating a weak model 100,000 times produces a very precise version of a weak assumption.

Injuries are still difficult to price

The most sophisticated rating can be undermined by information not fully captured in historical data: the player’s body.

Tennis injuries are rarely binary. A player can be healthy enough to compete and still be below normal level. A shoulder issue can reduce serve speed without forcing a retirement; a leg problem can become more costly after two hours; fatigue may be irrelevant in a routine first round and decisive after consecutive long matches.

Public information is uneven: some players discuss physical problems openly, others reveal little, and medical timeouts can mean different things.

For a model, the choice is uncomfortable. Ignore uncertain injury information and the forecast becomes stale. React too aggressively and it risks overfitting noise. A sensible system should therefore widen uncertainty rather than pretending that “fit” and “injured” are clean categories.

Head-to-head records need context

Head-to-head statistics remain popular because they are simple. Models have to treat them more carefully.

A 5-1 record can be meaningful if the matches were recent, played on similar courts and contested by players whose games have not changed dramatically. The same record can be almost useless if meetings came years ago or on different surfaces.

Most opponents simply do not meet often enough to create a large, stable sample. Giving head-to-head history too much weight can make a model worse. Yet ignoring it completely can miss genuine stylistic friction.

Good modeling asks whether the matchup history contains information beyond general player strength. Bad modeling simply counts wins.

What 2026 has taught us before the US Open

The three majors already completed this season have produced a useful mixture of confirmation and disruption.

The confirmation is easy to see. Alcaraz, Zverev and Sinner turned elite pre-tournament status into major titles. Their victories support the value of long-term ratings, surface-specific strength and the stabilizing effect of best-of-five tennis.

The disruption is just as important. Chwalinska’s run from qualifying to the Roland-Garros final showed how far a low-probability path can extend once a player survives the first upset. Noskova’s Wimbledon title showed how quickly an improving player can outrun a rating based partly on older information. Andreeva’s Paris victory showed that age itself is a weak variable unless it is connected to performance data.

None of those outcomes require us to choose between “data” and “intuition.” Models work best when they are allowed to be uncertain and when the person reading them understands what the number means.

The mistakes to watch for in New York

As the US Open approaches, the most revealing misses are unlikely to come from a model failing to identify a 100-to-1 champion. Those events will always exist.

More important are systematic mistakes. Does the model remain too loyal to ranking after a player’s level has clearly changed? Does it overreact to a small sample of summer hard-court matches? Does it ignore the quality of opponents behind a winning streak? Does it treat a recent injury as irrelevant because the player completed the previous match? Does it give head-to-head history more weight than current serve and return performance? Does it produce tournament probabilities before accounting properly for the draw?

Those are fixable errors. Randomness is not.

If a 75 percent favorite loses after the underdog plays the match of a career, the probability may still have been reasonable. If a model repeatedly labels obviously compromised players as heavy favorites because its injury information updates too slowly, that is a genuine weakness.

A model should be criticized for bad probabilities, not merely for surprising results.

The best forecast will still leave room for surprise

The appeal of the US Open has never depended on perfect predictability. New York forces every assumption to survive contact with heat, night sessions, crowd pressure, tiebreaks, physical recovery and opponents who refuse to follow the expected script.

Forecasting can organize that uncertainty better than intuition alone. It can quantify differences the rankings blur, adjust for surface and opposition, and prevent one dramatic recent result from dominating the analysis.

What it cannot do is remove the possibility that a qualifier catches fire, a young player improves faster than the rating system can react, or a championship favorite loses a few points that matter more than thousands that came before them.

That is the real scorecard for the models at the 2026 US Open. The question is not whether they can tell us the ending in advance. It is whether, before each match begins, they assign the uncertainty intelligently.

So far in 2026, the evidence points in both directions. The models have been right to keep trusting the sport’s strongest players. They have also been reminded that the gap between “unlikely” and “impossible” is where Grand Slam tennis often becomes most interesting.

← Back to Blog

Related Articles