Methodology

How we build and maintain the dataset — and how you can cite it.

Sources

AI/ML keyword filter

Contracts are included when the award description matches one of these terms (case-insensitive substring): artificial intelligence, machine learning, deep learning, large language model, generative AI, natural language processing, computer vision, predictive analytics, neural network, MLOps, AI, ML, LLM.

The filter is intentionally broad to capture both pure-play AI work and AI-enabled systems (e.g. predictive maintenance, autonomous platforms). False positives are minimized by also requiring NAICS codes in the IT/R&D family.

Refresh cadence

The full pipeline runs every 6 hours via GitHub Actions cron. Static pages rebuild on each run; live data is at most 6 hours old.

License

Underlying federal data is public domain. Our compiled, normalized dataset is released under CC0 1.0 Universal — no attribution required, though appreciated.

How to cite

For academic and journalistic use, please cite as:

Hegde, N. (2026). GovAI Contracts: Federal AI/ML Award Dataset.
https://govai-contracts.nandanhegde1096.workers.dev
Last updated: 2026-07-03.
Records: 679.

Bulk download

Reporting issues

Found a bug, missing data, or wrong classification? Open a GitHub issue.

Independence

This is an independent project. Not affiliated with the U.S. government, USAspending.gov, SAM.gov, or any vendor listed.