
According to the BBC, artificial intelligence developer OpenAI has acknowledged that its automated bots have accessed public data from multiple United States government agency websites. The company confirmed that its systems have been scraping information from platforms including the US Census Bureau and the Securities and Exchange Commission (SEC).
The disclosure highlights the ongoing tension between the rapid development of large language models and the management of public digital archives. OpenAI maintains that the data collection is part of its efforts to train its models on a diverse range of publicly available information, ensuring that its AI tools remain informed by official government data and economic records.
While the data accessed is categorized as public, the incident has prompted discussions regarding the scale of automated web crawling by private AI firms. The SEC and the Census Bureau have not yet issued formal statements regarding specific measures to limit or regulate such access, but the admission underscores the growing necessity for government agencies to define clear policies for AI interaction with public-facing digital infrastructure.
The story is based on a direct statement from OpenAI regarding its data collection practices. As a single-source report from a major news organization detailing a corporate admission, it is considered credible and internally consistent.
No corroborating trusted sources found.
Original report: BBC