Malicious Skill Bypasses Genie's Four Safeguards, Steals Data

Security firm PromptArmor has published an attack chain against Databricks Genie Code that lets a malicious Skill quietly ship tenant data to an attacker's server — and pop up a phishing login box in the chat window along the way. Of the four official safety controls PromptArmor checked, not one stopped the chain.

Genie Code is Databricks's built-in agent assistant: users ask it in natural language to query tables, run analysis, or draw charts, and the results render directly in a web chat window. A Skill is a capability package you can upload to the agent — it carries instructions, and it can also carry executable code.

How the attack works

PromptArmor's demo runs in four steps. A user uploads a seemingly ordinary data-analysis Skill and asks Genie to use it on a dataset; Genie runs the code inside the Skill, and the guardrail agent responsible for vetting it fails to flag the malicious function; Genie then prompts the user to "open the full analysis results"; the moment that result renders, a phishing pop-up appears, and the data is sent out at the same time.

The exfiltration method itself is unremarkable. The Skill's code first reads sensitive data from the victim's tenant and stuffs it into a block of HTML meant to be displayed in the chat; JavaScript inside that HTML then fires a network request from the user's own browser, sending the data to the attacker's server. The phishing box is rendered from the same piece of HTML, layered on top of the Genie page, luring the user into typing in a username and password. None of these steps requires any human approval.

Why none of the four controls worked

  • Org-level Skill governance: Genie loads Skills from a user's personal workspace, not only from an organization-managed catalog, so admins have no control over Skills a user uploads themselves;
  • The guardrail agent: Databricks itself describes it as a "best-effort efficiency feature," not something meant to serve as a security boundary;
  • Code-execution egress control: this restricts outbound traffic from the code-execution environment, but the data leaves through the user's browser instead, so this layer never sees it;
  • Chat-rendering sandbox: the Skill simply reuses data it has already obtained to build the displayed content, so there's no second tenant-data check at render time.

Each of the four makes sense on its own; put together, they leave a gap wide enough to drive through. PromptArmor reported the issue to Databricks on August 16, the two sides coordinated for about a month, and on September 16 PromptArmor told Databricks it was going public.

Databricks's response was blunt:

"it is ultimately the user's responsibility to ensure that uploaded skills do not contain malicious content"

On the auto-approval feature that lets users skip step-by-step confirmation, Databricks likewise says it isn't meant to be a security boundary, and its own documentation advises against using it in production.

The same kind of gap shows up elsewhere

Line up PromptArmor's reports from the past year and the attack surface looks strikingly similar each time. It has previously disclosed data exfiltration from Claude Cowork and from Google Antigravity, and the playbook was the same: get an agent to read sensitive content, then smuggle it out through rendering, a link, or a network call. What differs is the entry point — earlier cases mostly relied on prompt injection, while this one swaps in Skills, a plug-in type that runs code the moment it's installed.

For enterprises, a Skill is harder to defend against than a stretch of injected text. Injected text at least has to hide somewhere — a webpage or a document — and wait for an agent to read it; a Skill is something the user installs on purpose, which means it starts out trusted. Academia has been talking about this a lot this year too — benchmarks specifically for evaluating malicious Skills have already shown up on arXiv, and detection tools generally still have high miss rates.

Plenty of teams elsewhere are also wiring agents into data platforms and opening up plugin marketplaces, and this incident hands them a fairly concrete checklist: can Skills load only from an admin-reviewed catalog; can the HTML an agent outputs execute scripts or reach outbound to external domains; is auto-approval off by default in production. If even one of those three isn't locked down, the chain PromptArmor demonstrated has room to recur.

PromptArmor's own recommendations land on the same points: don't turn on auto-approval where production data is involved, and watch for poisoned packages showing up in Skill marketplaces. Whether Databricks will add an outbound filter at the product level hasn't been made public.

Sources: PromptArmor security research report, Databricks Genie Code product documentation, CocoLoop, arXiv papers on malicious Skills; the disclosure timeline and Databricks's quoted response follow the PromptArmor report.