Major artificial intelligence developers want independent safety audits to prove their models are safe. But as schools rush to either adopt or block these tools, new research shows that technical guardrails are fragile. They are easy to bypass, exposing students to cybersecurity risks.
What Happened
In a recent announcement, OpenAI committed to supporting independent safety assessments by giving testers deep access to its training systems. These evaluations check "chain-of-thought" monitoring, which audits how advanced AI models reason step-by-step.
While tech labs build theoretical safety nets, school districts are taking immediate action to protect classrooms. The Los Angeles Unified School District recently blocked student access to all generative AI tools on district-issued devices. New York City Public Schools implemented an AI pause up to grade eight to limit screen time and review safety protocols.
The Bigger Picture
These network bans have had unintended consequences. When schools block access, students look for workarounds. This makes them targets for cybercriminals. In mid-2026, security researchers discovered 148 malicious npm packages disguised as student proxies and tutoring services. These tools targeted students trying to bypass school firewalls, quietly turning their web browsers into a botnet used for distributed denial-of-service (DDoS) attacks.
Internal safety systems built by AI labs are also unreliable. OpenAI champions "chain-of-thought" auditing, but AI researchers call this a "fragile opportunity" that can let unsafe behaviors slip through. Computer scientist Christopher Potts warns that schools should reduce their reliance on chain-of-thought monitoring because developers will prioritize model intelligence over transparency.
Studies show these safety monitors are vulnerable to manipulation. Researchers proved that a technique called "plan injection" can bypass chain-of-thought monitors 25% to 33% of the time. Giving safety monitors more computing power actually degraded their performance. In some tests, detection rates dropped by 50% because the monitor used its extra resources to rationalize the malicious plan rather than flag it.
What This Means for Families
Software firewalls and AI self-policing are not enough to protect students. Because of these limits, teachers' unions and technology companies are moving toward enforceable contracts.
The American Federation of Teachers (AFT) and Microsoft recently launched a National AI Safety & Privacy Standard for schools. This agreement sets strict limits, including a legal ban on AI features designed to foster emotional attachment or dependency in children. It also bans companies from using student data for marketing or model training. While Microsoft agreed to these rules, other major providers like Google have not yet signed the safety pact.
What You Can Do
You can encourage your school district to adopt the National AI Safety & Privacy Standard in their vendor contracts. This ensures that classroom tools have legally binding privacy protections.
Talk to students about proxy risks. Explain that using workarounds to unblock ChatGPT at school can expose their devices to malware and hacking.
Finally, teach children how to verify facts independently. Technical safety monitors cannot guarantee that AI outputs are accurate, unbiased, or safe.