Codex Security Cloud: scans, duplicates, and fixes
OpenAI reported a 36% duplicate share in an internal security sprint. Here's how Codex Security Cloud handles scans, evidence, patches, and Daybreak Blue access.
OpenAI's security presenters said 36% of discoveries in an internal sprint were duplicates. Codex Security Cloud's focus on removing repeated findings makes more sense with that number in mind.
The presenters also reported a 1% patch rollback rate, with a separate verification step checking fixes. Simon Willison's session notes record both figures. These are OpenAI's internal results, not an independent benchmark or a promise for your repository.
The priorities are clear: investigate a possible vulnerability, cut repeated findings, and give you a fix you can assess. As of September 30, 2026, OpenAI lists cloud access on the announced plans, while the docs still call it a research preview.
Codex Security Cloud scans beyond the pull request
The Cloud FAQ describes a plugin that scans connected GitHub repositories, validates likely vulnerabilities, and returns findings with evidence and remediation guidance. You run it on the web or in the desktop app.
OpenAI's DevDay recap adds on-demand and scheduled repository scans, checks of new commits, duplicate removal, and proposed fixes prepared in the cloud. That work can continue while your laptop is closed.
Ordinary Code Review examines a change. Security Cloud can examine a repository and keep security context for later commits. Choose based on your question: did this change introduce a problem, or does this service contain a reachable vulnerability?
The product overview also distinguishes the Cloud plugin from the local Codex Security plugin. Each needs its own configuration; installing one doesn't set up the other.
Connect GitHub, then start a repository scan
The official setup guide gives this sequence:
- Find and enable Codex Security Cloud in the ChatGPT plugin marketplace.
- Connect GitHub and grant access to the repositories you want assessed.
- Select New scan, choose a repository, and select a compatible cloud environment.
- Keep Repository selected under What to scan, then start the scan.
- Inspect its progress and artifacts in Scans, then review individual issues in Findings.
If a repository is missing, check the connection and permissions. If the plugin is unavailable, a workspace administrator may need to check access. Those are setup problems, not evidence of how well the scanner assesses code.
For ongoing review, use the documented Commit changes workflow. It monitors new commits and can examine existing history; the settings let you choose the environment, adjust the history window, and pause monitoring. Follow the setup page for onboarding, since it spells out that path more clearly than the launch's scheduled-scan wording.
Start with a service whose ownership and expected behavior you know. You need that context to judge whether a reported security boundary actually exists in production.
Read the reproduction evidence, not just the finding
The FAQ separates analysis, scanning, validation, and remediation. During validation, the scanner tries to reproduce likely vulnerabilities in an isolated environment and attaches commands, logs, and other evidence.
A failed reproduction leaves a finding unvalidated, rather than proving the code safe. The environment may be missing a dependency, or the reproduction may be incomplete. A persuasive explanation still needs a reachable attack path.
For each finding, check:
- The input an attacker controls and the permissions the attacker needs.
- The path from that input to the affected code or privileged action.
- The evidence that reproduced the issue, or the specific gap that prevented verification.
- The production assumption that makes the reported impact possible.
An alarming description of code you can't reach under the stated conditions needs more scrutiny. OpenAI's scan guidance also recommends recording the target revision, scope, model, and reasoning effort when assessing your first scan.
Fewer duplicates should still leave enough evidence
Willison's live blog records the session's discussion of duplicate removal and the difficulty of assigning findings to owners. The 36% duplicate share came from OpenAI's internal sprint.
A scanner can reach the same vulnerable helper through several routes and produce several reports for one cause. You still need to know which report owns the repair and which affected paths need regression coverage.
Review what was grouped before celebrating a smaller dashboard. Similar symptoms can have different causes; different endpoints can share a vulnerable implementation. In your own trial, check that grouped findings keep enough evidence to tell those cases apart.
Your project's duplicate rate and the time your team saves are still things you'll need to measure.
Put the threat model in SECURITY.md or Cloud settings
Code rarely tells you every assumption about who can call a service or what sits in front of it. Write those assumptions into the threat model.
OpenAI's local scan documentation supports a root SECURITY.md for persistent security guidance and nested files for directory-specific policies. Use them for threat models, invariants, finding criteria, exclusions, and severity context. These files supply policy context, not executable instructions; build and validation commands belong in AGENTS.md.
For Cloud monitoring, edit the model under the repository's monitoring settings, as the threat-model guide describes. Changes guide future scans. Include entry points, trust boundaries, authentication assumptions, sensitive data paths, and review priorities.
For example, describe whether an upload endpoint accepts unauthenticated requests, which service checks ownership, and where uploaded files are processed. That example gives the scanner and reviewer more to work with than an instruction to find security bugs.
Update that context when deployment or authentication changes. A detailed finding can still rest on a stale premise.
Review the patch before creating a draft pull request
The session's verify-fix step challenged proposed repairs. Public detail about the reported 1% rollback rate's denominator and measurement period is limited, so you have little basis for comparing it with your own results.
The Cloud setup guide tells you to choose Fix with Codex when available, review the proposed patch, then select Create draft pull request. The FAQ says patches aren't applied automatically.
A repair should close the demonstrated path while preserving legitimate behavior. Ask for a regression test that fails before the fix and passes afterward, plus checks for permitted cases. Blocking every upload might stop an upload exploit, but it can also break the product.
The current CLI reference documents validate and patch, with no standalone verify-fix command. The session used that term for the verification approach. Check your installed tool's help before using it as command syntax.
Who gets Daybreak Blue, and what the open-source CLI needs
The launch recap lists Security Cloud for Pro, Business, Enterprise, and Edu on desktop and web. You get access to models offered through Daybreak Blue without a separate Daybreak application. That access is through Security Cloud, rather than unrestricted access to every Daybreak offering.
The current overview retains the research preview label alongside the broader plan rollout. Plan access doesn't mean the product has shed that label.
The CLI and TypeScript SDK ship in the public @openai/codex-security package. Willison reported the CLI as open source, and OpenAI links its public repository from the docs. You still need suitable access and authentication to run scans.
The CLI reference supports stored ChatGPT credentials or an API key, plus estimated cost controls. Scans use local operating-system permissions and don't pause for interactive approval. The tool being open source doesn't make model usage free or local scans offline.
Keep your static analysis and human review. OpenAI's FAQ says Codex Security complements SAST. Judge it by the evidence and useful fixes it adds to your workflow; a longer list of findings isn't enough.
FAQ
Is Codex Security Cloud available on Plus?
The DevDay availability statement lists Pro, Business, Enterprise, and Edu. It doesn't list Plus.
Does Codex Security Cloud fix vulnerabilities automatically?
It can propose patches. You review the patch before creating a draft pull request in the documented cloud workflow.
What is the Daybreak model in Codex Security Cloud?
OpenAI says you get access to models offered through Daybreak Blue. The announcement gives no universal model ID for every scan.
Does the 1% rollback rate guarantee reliable fixes?
No. OpenAI reported it from an internal result at the security session. It doesn't guarantee results for your repository.
Does Codex Security replace SAST?
No. OpenAI describes it as a complement to static analysis and manual security assessment.
Sources
- OpenAI: DevDay 2026 recap
- Simon Willison: DevDay security session notes
- OpenAI: Codex Security overview
- OpenAI: Security Cloud setup
- OpenAI: Security Cloud FAQ
- OpenAI: Improving the threat model
- OpenAI: Run a Codex Security scan
- OpenAI: Security CLI reference
- OpenAI: Codex Security public repository