
Focus
AI Safety, LLM Guardrails, Adolescent Wellbeing
Motivation
Child Protection, Responsible AI, Digital Policy
About the project
This paper investigates strategies businesses can implement to protect teenagers from the negative impacts of artificial intelligence on privacy, mental health and cognitive development. It reviews existing guardrails and regulations governing widely used large language models (ChatGPT, Gemini, Claude, Grok) and tests their robustness through an experimental red-teaming methodology, using structured prompts and follow-up questions that simulate realistic teenage interactions across academic integrity, privacy, manipulation and rule-circumvention scenarios. The study finds that model safeguards often degrade over extended, iteratively manipulated conversations rather than single exchanges, and that responsiveness and safety trade off against one another. It recommends age-adaptive guardrails, crisis-response and self-harm-prevention mechanisms, and adversarial detection that remains consistent across long conversations.
Check out more projects



