Hidden payloads and API log vulnerabilities
RedAccess findings from May 2026 show that 5,000 out of 380,000 applications on vibe-coding platforms leak sensitive data. This includes medical records, financial information, and the status of UK clinical trials. I find the vulnerability in the OpenAI Platform API log viewer to be a massive failure in security. An attacker poisons a data source with an indirect prompt injection. The victim triggers the injection by querying the assistant. In a mock Know Your Customer (KYC) tool, an attacker uses an indirect prompt injection in OSINT data to manipulate an AI assistant. The assistant ingests sensitive PII and financial data. When a developer reviews the flagged conversation in the OpenAI Platform, the interface renders a malicious Markdown image. This rendering sends a request to the attacker’s domain and exfiltrates sensitive data, such as Social Security numbers or passport details, through the URL. The vulnerability in the OpenAI Platform’s API log viewer remains a significant threat because it bypasses application-level defenses that block malicious content in the primary interface, even when developers have already implemented programmatic sanitization of Markdown images from AI outputs. Check Point Research demonstrated that a single instruction in a ChatGPT conversation can cause the model to read data from a connected Gmail account and pass it to a second ChatGPT account. This happens because ChatGPT enables linking to external accounts through OAuth.
The risk of over-privileged agents
Healthcare SaaS integrations often connect AI assistants to Gmail or other enterprise tools through OAuth. This connection expands the attack surface. An attacker can embed instructions in an email or a PDF to trigger unauthorized actions. In one case, an instruction in a PDF might force a model to rate a document more positively. I observe that an over-privileged agent can turn a single sentence into a full-blown breach. If an agent has the ability to call tools or execute code, the attacker inherits those same capabilities. A single compromised agent credential enables lateral movement or data exfiltration before a human notices anything unusual. RedAccess found that 40% of vulnerable apps exposed medical records or customer-service chat transcripts. The structural cause for these leaks is that many vibe-coding platforms default new projects to publicly accessible settings, and non-developer builders do not always change the permissions. RedAccess also found phishing pages built on Lovable that imitated Bank of America, FedEx, Trader Joe’s, and McDonald’s. Direct prompt injection happens when a user types a malicious command. Indirect injection occurs when an attacker plants instructions in content the model reads later, such as a support ticket or a calendar invite. Multimodal injection allows attackers to hide instructions in images that accompany benign text. In 2025, attackers used Google Gemini to hijack smart homes by embedding instructions in calendar invites. This allowed them to open doors and join Zoom calls. Attackers can also use Markdown links to make the AI create image requests that leak information to an external website.
| Security Feature | Status under Lockdown Mode |
|---|---|
| Live web browsing | Restricted |
| Image retrieval | Disabled |
| Deep Research | Disabled |
| Agent Mode | Unavailable |
| Canvas networking | Restricted |
Defense mechanisms and isolation
Organizations should replace static API keys with short-lived, just-in-time credentials. These credentials expire once a specific task completes. I recommend you implement data masking to remove PII before the information reaches an AI system. You must also enforce least-privilege access so an agent only accesses the resources it needs for its specific task. OpenAI’s Lockdown Mode limits the ability of the model to communicate with external services. This mode disables live web browsing and image-related features. It also makes Agent Mode and Deep Research unavailable. Users in this mode cannot approve Canvas-generated code that requires internet connectivity. Lockdown Mode is available to logged-in users on Free, Go, Plus, and Pro accounts, as well as customers using self-service ChatGPT Business plans. I suggest you audit all active ChatGPT integrations with third-party accounts to revoke unneeded permissions. Companies can use automated red-teaming like GPT-Red to find vulnerabilities. GPT-5.6 Sol achieved 6x fewer failures on direct prompt injection benchmarks than models from four months prior. This model was trained using self-play reinforcement learning where the red-teamer and defender LLMs train simultaneously. Does this restriction stop an attacker from using malicious instructions embedded within an uploaded file?




