# An NBA betting model whose README and requirements file describe two different projects

> A Python pipeline that scrapes team stats and sportsbook odds into two SQLite databases, trains moneyline and totals models, and prints win probabilities, expected value and optional Kelly sizing. The documentation and the dependency file disagree, the neural network path ships the older scripts, and no accuracy or return figure appears anywhere on the page.

**kyleskom/NBA-Machine-Learning-Sports-Betting** — NBA sports betting using machine learning

- Repository: https://github.com/kyleskom/NBA-Machine-Learning-Sports-Betting
- Stars: 1,732 · Forks: 572
- Language: Python
- License: not declared
- Published: 2026-09-13 · Updated: 2026-09-13 · Language: en
- Canonical page: https://hysenlabs.com/projects/kyleskom-nba-machine-learning-sports-betting

## The stated package list and the requirements file disagree in four places

The requirements section asks for Python 3.11 and names Tensorflow, XGBoost, NumPy, Pandas, Colorama, Tqdm, Requests and Scikit-learn. The pinned file behind it contains colorama at 0.4.6, pandas at 2.1.1, sbrscrape at 0.0.10, tensorflow at 2.14.0, xgboost at 2.0.0, tqdm at 4.66.1, flask at 3.0.0, scikit-learn at 1.3.1 and toml at 0.10.2. NumPy is not on that list at all, and neither is Requests, which is the library both scrapers need to talk to NBA endpoints and to the odds source. The file adds two things the prose never mentions, sbrscrape, which is the scraper the odds step depends on, and toml, which pairs with the config.toml sitting in the repository root that no section explains. One line is commented out, tensorflow-metal at 1.1.0, so an Apple silicon machine has to uncomment and re-pin it by hand. Install with the pinned file rather than the prose list and the difference will not surprise you.

## The neural network path ships the older scripts and the newer ones are elsewhere

One note under the training section explains an inconsistency the commands above it do not. The neural network training scripts in the repository are described as the original versions with hard-coded dataset and model paths: they train on dataset_2012-24_new and save into Models/ with timestamped names, and the page adds that if you want configurable flags or feature and scaler sidecars you should switch back to the newer NN scripts. Switch back to implies the newer pair exists somewhere other than the default branch. Meanwhile the gradient boosting and logistic regression scripts take their dataset as a flag, and the example commands pass --dataset dataset_2012-26 or dataset_2012-26_new with --trials, --splits and --calibration. So two of the three model families are configured from the command line and one is configured by editing source, and the neural net is also the only path that trains on a season range two years older than the others.

## Three dataset names and one calibration flag the network path never takes

Read the dataset names across the training block and the versioning shows through. The XGBoost moneyline and totals scripts are pointed at dataset_2012-26. Both logistic regression scripts are pointed at dataset_2012-26_new, a different file with the same season span. The neural network scripts, per the note, use dataset_2012-24_new, an older span and a third suffix. Nobody on the page explains what the _new suffix changes, and since the two families are trained on different files, a comparison between them is not apples to apples unless the datasets are the same. Calibration is configured the same uneven way. Every XGBoost and logistic regression command passes --calibration sigmoid, while the neural network commands take no such flag at all, which follows from their being hard-coded. The expected value number the script prints is derived from a model probability, so a model without a calibration step is exactly the one whose expected value should not be read literally.

## Odds arrive through one scraped site and a dependency at version 0.0.10

The pipeline keeps two databases. Get_Data pulls daily team stats from NBA endpoints into SQLite, and Get_Odds_Data pulls sportsbook odds and scores from SBR into a separate SQLite database. Create_Games then merges team stats, odds, scores and days rest into the training set. The odds half therefore depends on a third-party scraper, and that dependency is pinned at sbrscrape 0.0.10, a version number that reads like an early release rather than a maintained one. The consequence is structural rather than moral: your features are only as reproducible as someone else's page markup, and a change there shows up as a data problem rather than as a bug report. The ingest is also incremental by default. Both fetchers normally take only new dates in the current season, and the history from 2007-08 onward exists because someone ran the backfill flags, either for everything or for one season at a time with --backfill --season 2025-26.

## Seven book identifiers and one of them is scoped to a state

The odds flag takes a single book name, and the supported list is fanduel, draftkings, betmgm, pointsbet, caesars, wynn and bet_rivers_ny. Six are national brands and the seventh carries a state suffix, so the identifier encodes where that book operates rather than only who it is. That is a small design detail with a practical edge: the flag is a free string matched against this list, so a book that only takes bets in one state will silently produce no odds elsewhere, and the run falls back to asking. The fallback is documented and matters for interpretation. When -odds is provided the odds are fetched automatically; when it is omitted the script prompts for manual odds and totals, which means the same main.py produces both a scraped run and a hand-entered run with no marker in the output saying which one produced the expected values on screen. The manual path is not a debug convenience either, because it is the only way to price a line the scraper does not carry, and a hand-entered line is indistinguishable from a fetched one once it reaches Create_Games. Anyone comparing two runs of the model should record which books and which method produced the odds, because the merge step will not do it for them.

## The documented web command starts Flask with the debugger on

The web app is two commands from the Flask directory, and the second one carries a flag that matters:

```bash
cd Flask
flask --debug run
```

The debug flag starts the development server with its interactive debugger and reloader, which is the right choice for looking at outputs on a laptop and the wrong choice on anything reachable from a network, since that debugger can execute code in the server process. The README describes the app only as a way to browse outputs, and says nothing about binding it to a local port, putting it behind a reverse proxy or adding authentication. The repository root holds Data/ and Models/ next to src/ and Flask/, and the training scripts write timestamped model files into Models/, so the directory the app would serve is the same one the pipeline writes into. That is convenient locally and is the reason to keep the app off any shared interface.

## Release tags encode seasons and the licence file is missing

The three most recent releases are named after NBA seasons rather than after software versions alone: 2025-26.1.0 from 2026-01-09, 2025-26.0.0 from 2025-10-26, and 2024.25.1.0 from 2024-12-21. Read together they say the project versions itself per season, which is a sensible choice for a model tied to a league year, though the oldest of the three writes the season with a dot where the two newer ones use a dash, so the separator changed once. The default branch is master and the last recorded push is 2026-09-12. The licensing is the bigger gap: no licence file appears among the top-level entries, which are the gitignore, a Colab notebook, Data, Flask, Models, the README, Screenshots, Tests, config.toml, main.py, notes.txt, requirements.txt, scripts and src. Without a licence file the default copyright applies, so the code is readable and not reusable.

## Conclusion

Read this as a data pipeline and feature-construction exercise rather than as a betting system, because that is what the repository can support on its own evidence. There is no accuracy number, no backtest, no season-by-season record and no bankroll simulation anywhere on the page, and the expected value output is computed from the model's own probability against a scraped line, which only means anything if that probability is calibrated. Three checks before you trust a run. The neural network scripts on the default branch have hard-coded dataset paths and take no calibration flag, so the calibrated XGBoost and logistic regression paths and the uncalibrated neural net path are not comparable. The odds come from one scraped source through a dependency pinned at version 0.0.10, so the pipeline inherits someone else's page structure. And the repository has no licence file, so the code is not yours to reuse by default. Treat the Kelly output as arithmetic on an unverified probability rather than as advice about a bankroll.

## FAQ

### What does this NBA machine learning betting repository actually do?

It predicts NBA game winners and totals, over and under, from team statistics and sportsbook odds. Data is pulled from the 2007-08 season onward, matchup features are built from team stats, odds, scores and days rest, and XGBoost and neural network models produce win probabilities and totals outcomes. main.py then prints predictions, expected value and optional Kelly Criterion stake sizing.

### Which sportsbooks can it pull odds from?

Seven identifiers: fanduel, draftkings, betmgm, pointsbet, caesars, wynn and bet_rivers_ny, one of which is scoped to New York by its name. Odds are fetched automatically when the -odds flag is given, and if the flag is omitted the script prompts for manual odds and totals instead.

### What does the pinned requirements file contain?

colorama 0.4.6, pandas 2.1.1, sbrscrape 0.0.10, tensorflow 2.14.0, xgboost 2.0.0, tqdm 4.66.1, flask 3.0.0, scikit-learn 1.3.1 and toml 0.10.2, plus a commented-out tensorflow-metal line for Apple silicon. NumPy and Requests appear in the prose list of packages but not in the file, while sbrscrape and toml appear in the file and not in the prose.

### How does the project size a bet?

It computes expected value and, when the -kc flag is passed, shows the Kelly Criterion bankroll fraction. Both are optional outputs of main.py alongside the predictions, and the Kelly figure is derived from the model's own probability, which is why the calibration flag passed to the XGBoost and logistic regression training commands matters.

### Does this NBA betting repository carry a licence?

No licence file appears among the top-level entries, which are the gitignore, a Colab notebook, Data, Flask, Models, the README, Screenshots, Tests, config.toml, main.py, notes.txt, requirements.txt, scripts and src. The project metadata also records no licence, so the default copyright applies to the code.

## Sources

- [Issues](https://github.com/kyleskom/NBA-Machine-Learning-Sports-Betting/issues)
- [kyleskom/NBA-Machine-Learning-Sports-Betting on GitHub](https://github.com/kyleskom/NBA-Machine-Learning-Sports-Betting)
- [README](https://github.com/kyleskom/NBA-Machine-Learning-Sports-Betting/blob/master/README.md)
- [Releases](https://github.com/kyleskom/NBA-Machine-Learning-Sports-Betting/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kyleskom-nba-machine-learning-sports-betting
