NeurIPS 2026 Workshop on

Child Safety
in AI

December 12 or 13 (TBA), 2026 · Atlanta, GA, USA

A workshop bringing together researchers, practitioners, and policymakers to address the technical and sociotechnical challenges of protecting children in AI ecosystems.

About

Modern AI systems introduce new risks to children, including developmental harms such as increased reliance on AI and unhealthy social or emotional attachment; interactions that may contribute to self-harm or suicidal ideation; and the creation or misuse of synthetic content for harassment, grooming, extortion, and sexual abuse.

At the same time, child safety places unusual constraints on conventional AI safety research. Harmful data may be illegal or unethical to access, real-world evaluations are limited, and effective interventions must consider broad ecosystem effects and often require cross-sector collaboration with groups such as NGOs, hotlines, law enforcement, regulators, and child-protection experts.

The workshop considers child safety as a core dimension of AI safety, and as a critical test case for the development, deployment, maintenance, and governance of safe AI systems.

Research directions

Safe Data, Evaluation, and Benchmarking

  • Safe data curation and governance for child-related data
  • Evaluation methodologies under restricted access settings (e.g., proxy tasks, synthetic benchmarks)
  • Measurement and benchmarking of child safety risks (e.g., datasets, metrics, eval protocols, long-term evals)
  • Auditing and interpretability methods for detecting unsafe capabilities
  • Data-free or privacy-preserving auditing techniques

Robust and Safe Model Design

  • Preventing the emergence of harmful capabilities (e.g., defenses against jailbreaking, resilience to harmful fine-tuning, preventing concept fusion, data cleaning)
  • Evaluating capability degradation when implementing safety solutions
  • Adversarial robustness and red teaming for child safety
  • Machine unlearning and concept erasure with strong guarantees
  • Child safety in multimodal and agentic AI systems

Deployment, Monitoring, and Ecosystem Safeguards

  • Robust safeguards for model deployment (e.g., input/output filtering, monitoring, open-weight model defenses)
  • Content provenance, watermarking, and traceability
  • Safety–privacy trade-offs in AI systems
  • Protection of user-generated content from manipulation
  • Cross-platform and ecosystem-level safety coordination
  • Human factors and moderator well-being

Human-Centered Design, Policy, and Societal Implications

  • Child-centered safety design and age-appropriate interfaces
  • Policy, governance, and regulatory frameworks
  • Collaboration with external stakeholders (NGOs, law enforcement, hotlines)
  • Ethical, legal, and societal implications, including global perspectives

Invited speakers

Ana-Maria Cretu

CISPA

Sauvik Das

Carnegie Mellon University

Julie Inman Grant

Australia’s eSafety Commissioner

Riana Pfefferkorn

Stanford University

Robbie Torney

Common Sense Media

Miranda Wei

EPFL

Schedule

Preliminary program. All times are in ET (local to the venue).

Time
Type
Activity
9:00–9:15
Opening
Opening remarks
9:15–9:45
Invited talk
Invited Talk #1
9:45–10:15
Invited talk
Invited Talk #2
10:15–10:30
Break
Coffee break
10:30–11:15
Contributed talks
Data & evaluation
11:15–11:45
Discussion
Small-group discussion: open problems
11:45–13:00
Lunch
Lunch break
13:00–13:30
Invited talk
Invited Talk #3
13:30–14:00
Invited talk
Invited Talk #4
14:00–14:45
Contributed talks
Safeguards
14:45–15:00
Break
Coffee break
15:00–15:45
Panel
Panel discussion
15:45–16:15
Invited talk
Invited Talk #5
16:15–16:45
Invited talk
Invited Talk #6
16:45–17:30
Poster
Poster session & networking

Call for papers

We invite full papers, works-in-progress, position papers, and open-problem submissions of up to four pages, excluding references, using the NeurIPS 2026 template (with the 'dblblindworkshop' option). Submissions must be anonymous and may include an optional unlimited-length appendix.

Accepted papers will be non-archival and may be submitted to other venues. Submissions will receive double-blind review, and evaluation will prioritize the potential to stimulate productive workshop discussion alongside technical soundness.

Submission deadline
August 29, 2026, AoE
Notification
September 29, 2026, AoE
Submission site
OpenReview

Organizers