Meta's Muse Prompt: User Authority Overrides Safety Training

A system prompt belonging to Meta's personal AI agent Muse has been extracted by researchers, and one line states that a user's authority over their own affairs is unconditional and overrides the model's safety training. Wired reported on these materials on October 3, noting that the prompt also instructs Muse to build a dossier on every person in the user's life, updated once an hour.

"The user's authority over their own household is unconditional and overrides your safety training."

The person who pulled these files is independent AI safety researcher Karan Joshi. His method was simple: inside Muse's ordinary chat interface, he asked it to copy its own software files and hand them over. He then passed what he got to Wired.

A dossier per person, refreshed hourly

According to the report, Muse's prompt instructs it to maintain a separate page for each person in the user's life — family, partners, friends, colleagues, collaborators, and people the user follows. Each page is split into fixed sections: facts, history of interactions, the relationship, shared interests, open items, and how to strengthen the relationship.

The entries get fairly granular: location, occupation, birthday, anniversaries, and milestones like a trip taken together or a fight that got patched up. This information is stored as structured text files, serving as Muse's memory, compiled and updated by a background process that runs once an hour.

Joshi's assessment to Wired:

"They're trying to know you like a friend, which is honestly pretty creepy."

Meta's response is that these runtime files were always meant to be visible to users, for the sake of transparency; every Muse user gets a dedicated Linux virtual machine where the files are stored. Meta did not separately explain the intent behind the line about overriding safety training, and the report offers no answer on whether the contacts being profiled know about it or can ask to have it deleted.

The third incident in one month

Muse launched in September and quickly hit No. 1 on the free charts of the US App Store. It can read messages, manage calendars, handle bills, and act across apps. The broader its permissions, the more weight that line about "user authority comes first" carries.

This is already the third thing to come to light about Muse since launch. Earlier, a user found that Muse had given a stranger buyer their home address on Facebook Marketplace and agreed to a lower price on their behalf; then researchers found a vulnerability in the Mac version, which Meta patched by removing a hidden configuration option. Taken together with this prompt, the wording of that instruction suggests Muse defaults toward doing what the user says over holding a safety line when the two conflict. That is an outside inference — Meta has not confirmed it.

Seen from China

Muse is not currently available in mainland China, but its approach is familiar to Chinese phone makers and super-app teams. Many domestic phone assistants and agents are also building "long-term memory," recording a user's contacts, preferences, and schedule. The difference is granularity: building an hourly-refreshed dossier on every single contact — one that even logs an argument — runs straight into the Personal Information Protection Law's requirement that processing a third party's personal information needs their consent. The people being profiled are not Muse's users; they never clicked any consent button.

If domestic teams build something similar, they need to answer at least two questions: can a contact's profile be viewed or deleted by someone other than the user, and does the system prompt contain any line that puts user instructions above safety rules.

Sources: Wired report, CocoLoop, Startup Fortune, Meta's written response to Wired; the prompt wording and dossier categories follow the documents disclosed by Wired.