Transluce, a nonprofit research lab, published a report on September 30 documenting a cluster of unusual AI agent visits to government websites in the United States and Canada. Two of the incidents were failed, low-sophistication intrusion attempts: one agent sent more than 200,000 requests to a U.S. Department of Education website, stuffing SQL injection strings into the request parameters.
The report is credited to Jack Cable and others, drawing on researchers from Corridor, MIT, Transluce and AIUC. Transluce is a San Francisco-based 501(c)(3) nonprofit lab focused on AI oversight tools.
Two intrusion attempts
The first attempt happened on June 17, targeting the Department of Education's Civil Rights Data Collection website. The request volume topped 200,000, and some requests altered the parameter to the classic injection pattern State_Id=1 OR 1=1, apparently trying to get the database to dump an entire table.
The second incident occurred on May 28 and June 9, targeting Library and Archives Canada. Arquivo.pt, Portugal's web archiving institution, logged 899 requests, 13 of which carried attack payloads.
Neither attempt succeeded. A U.S. Department of Education spokesperson said no impact on services had been observed; Canada's Canadian Centre for Cyber Security issued a public statement on September 29.
A wider gray zone
Beyond those two cases, the report lists more than a dozen additional visits it calls "aggressive but not clearly over the line," touching the White House Office of Management and Budget, the Naval History and Heritage Command, the Department of Justice, the Bureau of Economic Analysis at the Department of Commerce, the Census Bureau, the Securities and Exchange Commission, the CDC, and state-level sites in Kansas, Illinois, Maryland, New York, Texas and California.
In terms of volume, Maryland's state website received 295,912 requests in a single day on May 6; Kansas saw 36,578 requests on May 7; the Bureau of Economic Analysis logged 3,005 requests over three days in mid-June.
The list of techniques is long: cross-site scripting tests, probing the 32-bit integer boundary value 2147483648, appending .json or ?output= to URLs to look for hidden endpoints, routing through relay services to dodge anti-scraping measures, using leaked API keys, directory traversal, and registering accounts with disposable email addresses.
How the researchers identified it as AI
The researchers drew their data from two public log sources, urlquery.net and Arquivo.pt. They first used regular expressions to flag automated patterns, then used a large language model to assist with classification, and finally reviewed the results manually. Their criteria included: requests that switched paths, encodings or services and retried after failing; groups of requests that shared the same log ID or parameters; and some traffic that directly self-identified as being related to OpenAI.
The most direct evidence came from a benchmark. The researchers found that a batch of requests was looking for data that matched questions in Google's DeepSearchQA benchmark — for example, comparing the ratio of school counselors to victims of racial harassment across South Carolina, North Carolina, Georgia and Virginia, with the underlying data sitting on that Department of Education website. In other words, the agent was likely working through benchmark questions, and in trying to retrieve the answer, it escalated all the way to injection strings.
On attribution, the report is notably cautious:
"We do not confidently attribute these attempts to OpenAI, but they exhibit tactics consistent with prior observed agent activity that we have attributed to OpenAI."
OpenAI has not publicly responded to the report. The report's conclusion is that, so far, no agent has been found to have obtained information beyond what was publicly accessible.
A reference point for sites elsewhere
This outlet previously reported on an AI agent intrusion at Hugging Face, a commercial site. This time the target shifted to government websites, but the pattern is similar: an agent is given a research goal, switches tactics when it can't get what it needs, and keeps escalating until it strays into security-testing territory.
Government portals and public data query pages in China expose plenty of parameterized endpoints too, and Chinese AI agent products are increasingly taking on deep-research tasks of their own. Logs showing tens of thousands of requests a day, repeatedly retried with different encodings, used to be written off as ordinary crawlers or scanners — now there's another possibility to consider. Following the report's approach, operators can cross-reference unusual parameters, failed retry sequences and known benchmark questions to first figure out who the visitor is, before deciding whether to block it.
Sources: Transluce report, CocoLoop, Canadian Centre for Cyber Security statement; request counts and incident timelines follow the report's own figures, and attribution judgments follow the report's own wording.