Onset

OpenAI Previews Safety Checks With No Staff Access to Data

OpenAI Previews Safety Checks With No Staff Access to Data

The short version

  • OpenAI is testing automated risk monitoring that keeps eligible customers’ prompts and responses in their encrypted infrastructure.
  • The company receives limited alerts, while customers investigate their own records and choose whether to share the underlying content.
  • It remains an unvalidated preview, excludes consumer ChatGPT, and allows retention or human review for severe risks.
A blue, green, and gold gradient title card with a fine grid and white dots. Large centered white text reads: “Offering Zero Data Retention for frontier models.”

Latest update

Anthropic plans customer-cloud storage for 30-day safety logs

The option will cover Mythos and Fable enterprise customers, keeping the required logs in customer-controlled clouds rather than Anthropic’s cloud.

View 1 earlier update

OpenAI plans to start rolling out its hosted Private Safety Processing option in September

The customer-key-encrypted option is now being tested with early customers.

Full story

OpenAI has previewed a system designed to monitor how its most advanced AI models are used without giving company personnel access to customers’ prompts or responses.

The approach, called Private Safety Processing, is intended for eligible business and API customers approved for Zero Data Retention. In those deployments, OpenAI says interaction history remains in customer-controlled, encrypted infrastructure, where automated systems can look for risky patterns across a sequence of requests. OpenAI receives only a limited alert describing the category and severity of a suspected problem.

OpenAI’s proposed flow for reviewing API activity without giving its personnel access to customer content
via x.com

Customers would investigate alerts using their own records and decide whether to share the underlying material with OpenAI. The company retains exceptions for severe risks that may permit retention and human review, as well as documented procedures for images suspected of containing child sexual abuse material.

The system is an early-customer preview, not a generally available or independently validated product. OpenAI said a rollout and technical white paper are planned for September; it has not yet published detection results, false-positive rates, latency, costs or audit evidence.

The plan marks a different response to an industry-wide tension between privacy and monitoring. Anthropic requires 30 days of retention for its Fable 5 and Mythos 5 models, while OpenAI is betting that monitoring across requests can work without the AI provider keeping a readable archive. The preview does not apply to consumer ChatGPT subscriptions, and Zero Data Retention remains approval-gated and limited to certain API endpoints.