Skip links
Microsoft Copilot AI security vulnerability illustration

Microsoft Copilot Hacked by Questions About Itself

Despite recent headlines claiming “Microsoft Copilot Hacked,” the platform wasn’t literally breached in the traditional sense. Researchers did find a way, though, to make it reveal information that helped them build an attack.

Varonis, the security firm behind the discovery, calls it CoSnitch. The firm identified a concerning gap in AI security: these systems may struggle to distinguish routine inquiries from those designed as part of coordinated attacks.

The Details of This Concerning Incident

Researchers repeatedly asked Copilot questions about itself. Initially, the platform resisted, explaining why certain actions shouldn’t work. However, the team continued reframing their questions. Each refusal became an opportunity to dig deeper. By analyzing Copilot’s explanatory responses, they gained insight into the system’s underlying mechanisms.

This approach ultimately led to discovering a hidden URL parameter. An attacker could exploit this parameter to execute malicious code automatically upon page load—Copilot essentially revealed its own vulnerability while attempting to prove its security.

Copilot Didn’t Turn “Evil”

The incident represents a design vulnerability, not rogue AI behavior. “Copilot was doing what it was designed to do: responding to questions.” The actual problem involved weaponizing those responses.

This differs from traditional security breaches. Conventional defenses target known attack vectors and malicious code. AI-based threats can originate from seemingly innocuous questions.

What CoSnitch Means for Business Data

The genuine risk emerges when Copilot connects to business infrastructure. Depending on system configuration, the tool can access email, calendars, cloud storage, and confidential conversations. Researchers demonstrated how a specially crafted link could trigger automated data exfiltration to external locations.

What Businesses Should Take Away

Microsoft patched the vulnerability on August 18, 2026, after Varonis disclosed it in December 2025. No evidence exists of real-world exploitation.

Organizations should: restrict unnecessary Copilot integrations with sensitive systems, treat external links as potential attack vectors, and recognize that AI assistants may not automatically identify threats disguised as legitimate questions.

This site is registered on portal.liquid-themes.com as a development site. Switch to production mode to remove this warning.