Skip to the page
History

AI and Machine Learning in Sports Betting Models

Everyone and their mother is running an AI betting model now. Does it work? Mostly no. Sometimes yes. Here are ten takes on where the actual edge is and where the snake oil lives.

By Priya Desai6 min read
an AI model training dashboard with sports betting prediction outputs and machine learning charts

AI and machine learning in sports betting. Okay. The sportsbooks are using it. The sharps are using it. Now your cousin who just discovered Python is telling you about his XGBoost model and the 58 percent win rate he got on his backtests. Here is the rundown.

1. The sportsbooks have better models than you

This is the starting point. Every major sportsbook runs machine learning on pricing. Pinnacle has for years. DraftKings, FanDuel, Caesars, and every significant US and European book has invested heavily in algorithmic pricing. Their models consume real-time injury data, weather, lineup changes, money flow from sharp accounts, everything.

Your model, trained on public historical data from a Kaggle dataset, is not going to beat their model on average. You need an information edge they do not have, or a specific niche they are not focused on.

2. Backtest numbers lie

I cannot say this enough. When your cousin tells you he got 58 percent on backtest, ask him if he accounted for market efficiency at the time of the bet. He probably used closing lines. Closing lines are contaminated by information the bettor would not have had at the time of placement. A model that looks profitable against closing lines will often be break-even or losing against opening or mid-market lines that you would actually bet into.

The correct backtest uses lines available at the time of decision, not after the fact. Most public backtests do not do this properly. Cousin's 58 percent is probably 51 percent in real-time reality, and 51 percent is break-even against minus 110 vig.

3. Feature engineering matters more than the algorithm

Everyone is obsessed with XGBoost versus random forests versus neural networks. That is not the interesting question. The interesting question is: what features are you feeding the model?

Public box score stats are available to everyone. The sportsbook has them. The edge is in features that are not in the standard public data: detailed possession-level tracking data, biomechanical injury indicators, sleep and travel patterns, referee-specific fouls tendencies, lineup interaction effects. The algorithm is a commodity. The features are where the money is.

4. Overfitting is almost guaranteed

Sports betting has a small number of events compared to most machine learning applications. An NBA season is 1,230 games. A full NFL regular season is 272. Even if you have ten years of data, you are looking at 10,000 to 12,000 NBA games and 2,700 NFL games. A modern ML model with thousands of features will overfit this like crazy unless you are extremely disciplined about cross-validation and regularization.

The signs of overfitting: great backtest results that do not replicate in live betting, a model that needs constant re-training to stay performant, predictions that are wildly confident on edge cases. You know what I mean? If the model says the Lakers at minus 3 is 62 percent, and the Lakers are actually historically 58 percent in that situation, you have a problem.

5. Injury data is the real frontier

The single most exploitable edge for a retail bettor with a model is injury data. The sportsbook's injury model is fast but not perfect. A bettor who pays attention to beat reporters, practice reports, and specific injury terminology (day-to-day versus out versus questionable) can sometimes price an injury impact more accurately than the book.

Combining injury data with lineup projections and expected minutes is where several of the genuinely profitable small-scale models I am aware of get their edge. It is not sexy. It is scraping and parsing. It is the work most modelers do not want to do.

6. Market-based models can beat fundamental models

This is subtle. Instead of trying to predict game outcomes from fundamentals, some profitable models predict how the betting market will move from opener to closer. If you can forecast line movement, you can bet the opener in the direction the line will move, and you are capturing closing line value mechanically.

The books know this and try to price openers to minimize it, but retail-facing books (DraftKings, FanDuel) are more willing to post softer openers than sharp books (Pinnacle, Circa). The sharper books get the fair market price faster, which means fading them into softer retail books can produce positive EV even without a fundamental model.

7. The books will limit you

How many times am I going to say this. If your model actually wins, the books will limit you. Your betting limits will drop. You will be asked to stop betting certain markets. In extreme cases, your account will be closed.

This is the real scaling problem in AI betting. Even if you have a profitable model, you can only bet so much before the book restricts your action. Building a profitable model is hard. Deploying it at scale without getting limited is harder. Some syndicates run dozens of accounts through friends and family to get around this. That is a whole other can of worms and has its own legal and relationship issues.

8. Kelly sizing and variance

Even a good model produces a thin edge. Two percent is a great edge in sports betting. Three percent is elite. At those edges, you are looking at Kelly bet sizes of 1 to 4 percent of bankroll per bet, and you need hundreds or thousands of bets before your results converge on the expected value.

A year of betting 500 games with a 2 percent edge and quarter-Kelly sizing produces expected ROI around 3 to 5 percent with a standard deviation that can produce a flat or losing year well within one standard deviation. You need bankroll discipline to survive.

9. Live betting models are hardest

As I discussed in the micro markets piece, live in-game betting has the thinnest pricing margins and the fastest movement. A live-betting model has to update predictions several times per second, has to handle latency, has to get past sportsbook bet-blocking for winners. Very few people are doing this profitably at retail. Institutional players are, but retail, mostly no.

10. The human in the loop beats full automation

The working profitable models I know about combine ML output with human judgment. The model identifies spots where it thinks there is an edge, a human reviews the spot, applies qualitative overrides (is the coach going to rest starters? is there a weather issue the model is underweighting? is there breaking news?), and bets the filtered subset.

Fully automated betting, where the model places bets without human review, tends to underperform human-in-the-loop systems by a meaningful margin. The model misses context. Humans miss math. Together, they beat either alone.

All right. That is AI in sports betting. Short version: hard, possible, rarely profitable at retail scale. If you are doing it for fun and learning, knock yourself out. If you are expecting to make real money, be very honest with yourself about the probability of it working. Most people who try, do not. That is just the math of small edges in a liquid market with sophisticated counterparties.