Safety transparency
This page describes system behavior without disclosing thresholds that would help bypass controls.
What is checked
Text input is reviewed before Provider use, and Provider output is reviewed before browser delivery. Existing approved conversations remain private to the account. Image upload is disabled.
What is recorded
We record keyed content fingerprints, categories, detector and policy versions, decision actions, timestamps, related account and model identifiers, appeal results, and audited reviewer actions. Raw content is not copied into ordinary logs.
Operational metrics
The safety team reviews block rates, detector disagreement, review time, repeated abuse, appeal overturns, service availability, and evidence-access events. Metrics are used to tune policy and detect failure—not to claim perfect accuracy.
Limits
Automated classifiers can be wrong and cannot determine copyright ownership, permission, fair use, or criminal liability. Serious consequences require human and, where appropriate, legal review.