Attacker stole a METR API key, used $600K worth of credits, and no one noticed for weeks
Security
The model provider gave METR the credits for free. An actual customer would not have been so lucky
AI model testing organization METR has disclosed two attacks that happened earlier this year, including one in which an attacker stole an API key and spent three weeks consuming public-model credits worth about $600,000.
METR (short for Model Evaluation and Threat Research) found no evidence that the attackers accessed sensitive information in either incident, and the org said it investigated both with security experts.
METR researchers worked with OpenAI to investigate how its agents hacked Hugging Face, and on Monday, it disclosed two of its own security snafus.
“In March 2026, attackers stole an API key for inference on public models and consumed a substantial amount of credits,” the nonprofit disclosed in a Monday report. “In May 2026, we observed attackers systematically probing our publicly accessible infrastructure, including an unsuccessful attempt to access internal data via an inadvertently exposed endpoint.”
From fail-open bug to model-credit theft
The March incident involved a METR researcher who didn’t have access to sensitive information - including model data and credentials, as well as information about model architectures, training, and release dates. The researcher used agents running on a personal EC2 instance that was “intentionally” left publicly accessible behind Google authentication. The instance contained an API key for METR’s public models account.
According to METR’s account, a “vibe-coded app” included a fail-open bug that disabled authentication, and this exposed the system to the public internet for several days.
“We suspect that the attacker found the instance by looking through recently-registered websites (e.g. in certificate transparency lists) to find vibe-coded sites with high-signal keywords relating to LLMs or agents, for purposes of harvesting potentially exposed model provider API keys,” the AI research org wrote.
Once the attacker found the app, they prompted an agent to reveal its model provider API key, then added an SSH key to maintain persistent access, and over the next three weeks used the stolen credentials to consume API credits on public models worth about $600,000. Luckily for METR, the unnamed model developer had given the credits to the nonprofit for free.
How do you not notice the 'large illicit usage?'
METR does answer the question on everyone’s mind in the report: Why its researchers didn’t notice the “large illicit usage?”
There are several reasons for this. First, the model testing operation regularly runs evaluations that use a lot of tokens, and this means the organization is “very acclimated to getting lots of weird rate limit and API errors.” So the high usage didn’t look that out of the ordinary.
Plus, since the tokens were free, METR didn’t accrue a large bill, and at the time there was no way to put a spending limit on keys like the one that was stolen.
In response to the March incident, METR says it improved its security infrastructure, protocols, and review process, and will continue to invest in security. To this end, it also hired a security lead, and plans to add more security staff.
Crims used agents to try to access frontier models
The second incident happened in early May, when “METR became the target of a sustained external attack campaign.”
After being “tipped off” that attackers who appeared financially motivated may have been trying to gain illicit access to frontier models, METR watched the intruders probe its publicly accessible infrastructure. They also used agents to find ways to gain initial access, including automated vulnerability discovery, credential stuffing against authentication providers, attempting OAuth token grants, scanning newly deployed services, and phishing attempts.
At the same time, METR unintentionally “exposed a read-only SQL query mechanism via our public transcript viewer.” While queries were scoped to public data by default, a bug allowed access to unpublished evaluation data, and “some sensitive model data was accidentally included in this database.” However, there’s no evidence that the attacker found the exploit or accessed any non-public data, according to the model testing body.
An independent bug hunter discovered the vulnerability and reported it to METR, which paid the researcher a bounty, and took the API offline.
In response, METR says it now uses an isolated production environment for public-facing applications that is separate from its internal infrastructure.®
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Wow
0
Sad
0
Angry
0
Comments (0)