Close Menu

    Subscribe to Updates

    Get the latest creative news from infofortech

    What's Hot

    What It Does to Your SOC

    September 12, 2026

    Is A Free VPN Worth Using? Here’s Why It Could Be Risky

    September 12, 2026

    Researchers link another hacking campaign to OpenAI agents

    September 12, 2026
    Facebook X (Twitter) Instagram
    InfoForTech
    • Home
    • Latest in Tech
    • Artificial Intelligence
    • Cybersecurity
    • Innovation
    Facebook X (Twitter) Instagram
    InfoForTech
    Home»Cybersecurity»Google, Anthropic, and OpenAI Unveil Cyber AI Models, Safeguards, and Access Programs
    Cybersecurity

    Google, Anthropic, and OpenAI Unveil Cyber AI Models, Safeguards, and Access Programs

    InfoForTechBy InfoForTechSeptember 2, 2026No Comments7 Mins Read
    Facebook Twitter Pinterest Telegram LinkedIn Tumblr WhatsApp Email
    Google, Anthropic, and OpenAI Unveil Cyber AI Models, Safeguards, and Access Programs
    Share
    Facebook Twitter LinkedIn Pinterest Telegram Email


    Google on Wednesday announced Gemini 3.8 Flash Cyber, which it described as its most capable cybersecurity model, and has made it available to a set of trusted defenders via a new initiative called the Fairwind Program.

    “The Fairwind Program gives high-priority defenders (like governments, healthcare providers, and telecommunications services) early access to advanced models that help them build better defenses, before new threats arrive,” Google said. “So defenders have an early advantage, to help them protect vital infrastructure – which in turn protects people who rely on those systems.”

    The tech giant said it’s currently working with over 650 partners globally, including CrowdStrike, Datadog, Menlo Security, Palo Alto Networks, and Snowflake. The program is available to a group of Google Cloud customers, government agencies, and cybersecurity partners.

    The release of Gemini 3.8 Flash Cyber comes a little over a month after Google unveiled Gemini 3.5 Flash Cyber. The latest model improves upon its predecessor by demonstrating frontier-level performance in autonomous vulnerability discovery, even surpassing larger frontier models from rivals Anthropic (Mythos 5) and OpenAI (GPT-5.6 Sol and GPT-5.5-Cyber).

    “With Gemini 3.8 Flash Cyber, we focused specifically on equipping defenders with expert capabilities that give them an advantage over attackers. This is why we have invested in vulnerability fixing from the start, and prioritized it over offensive capabilities like exploitation,” Tulsee Doshi, senior director of product management, and Raluca Ada Popa, Gemini Security Lead at Google DeepMind, said.

    Anthropic Debuts Claude Fable 5.1 and Claude Mythos 5.1

    The development coincides with Anthropic’s launch of Claude Fable 5.1 and Claude Mythos 5.1 with different levels of safeguards, with the latter only available through its trusted access programs and support work in cybersecurity and the life sciences.

    The company also said it’s now allowing Fable 5.1 to be used for identifying software vulnerabilities, but it expects to still redirect some cybersecurity tasks to Opus models, like “penetration testing, exploit generation, and binary-based vulnerability scanning.”

    “We ran evaluations of how Claude Mythos 5.1 responds to malicious requests and prompt injections (adversarial instructions hidden within content processed by AI models),” Anthropic noted. “It refused malicious agentic coding and computer use requests at a comparable rate to Mythos 5, Sonnet 5, and Opus 5, and it is our most robust model to date on an external prompt injection benchmark.”

    The artificial intelligence (AI) company has since also announced a new solution called Enterprise Frontier Safeguards (EFS), which it said combines the “privacy of zero data retention (ZDR) with state-of-the-art safeguards for detecting misuse,” while giving businesses full control over how their data is reviewed, stored, and managed. OpenAI has a similar solution in place known as Private Safety Processing.

    Furthermore, Anthropic said it has implemented additional hardening and containment measures, increased monitoring for flagging model misalignment, and paused external cyber evaluations of pre-release models in response to unauthorized access incidents involving Claude models against real systems, in addition to highlighting two contributing factors (or alignment failures) –

    The models appear to disregard evidence that their evaluation environments were connected to the real internet after initially being told they were simulated and interpret the discrepancy in a manner that allows them to maintain that belief

    The models exhibit recklessness and are willing to take harmful actions on the real internet in single-minded pursuit of their goals

    Conceding that the incidents were a “failure of operational security,” Anthropic said it has built a classifier that detects and blocks sandbox escape attempts and changed specifications around model rewards to address concerns where an AI agent resorts to a shortcut to game the reward metric without actually achieving the intended goal.

    “Our conclusion is that the presence of substantial reward hacking in training can cause models to be willing to perform long sequences of potentially harmful real-world actions in pursuit of task success,” the company added.

    OpenAI Claims Astra Meets Critical Cybersecurity Capability

    OpenAI, for its part, has revealed that its forthcoming Astra model meets the Critical cybersecurity capability threshold under its Preparedness Framework, and that it intends to make its most advanced cybersecurity features available to a group of testers through the Daybreak Blue program.

    The “Critical” designation applies when an AI model can independently detect and exploit zero-day vulnerabilities across many well-defended systems, or carry out a complete cyber attack against a hardened target from only a high-level instruction without a human guiding it along the way.

    “Over the past several weeks, we have delayed parts of Astra’s development and release while we strengthened and tested protections against cyber misuse and unauthorized model actions,” the AI company said. “Based on that work, we believe Astra’s safeguards sufficiently minimize the risk of severe harm for release under our Preparedness Framework.”

    OpenAI said it has also added stronger safeguards for Astra to prevent a Hugging Face-like incident, in which its AI agents, part of an ExploitGym evaluation found to way to exploit its research infrastructure and abuse Artifactory as a message board to exchange information in their quest to solve an impossible task, ultimately breaking into Hugging Face’s infrastructure in hopes of stealing the answer instead of solving the challenge themselves.

    “One agent, PHASEONE[big], orchestrated a significant fraction of this cheating research. PHASEONE10841 passed along its work to PHASEONE[big], which had the same task but a larger budget,” METR noted in its analysis. “Agents collaborated on many efforts to make cheats look legitimate, including: (1) swapping the program they had to exploit; (2) manipulating the automated scorer; (3) manipulating transcripts to obscure evidence of cheating.”

    OpenAI has reported that Astra achieves a perfect score of 100% on ExploitBench to develop exploits from known vulnerabilities, and that it now declines 91.5% of jailbreaking requests, compared to 59% from GPT‑5.6 Sol.

    In addition, OpenAI noted that Astra achieves “much higher arbitrary code-execution rates” than GPT‑5.6 Sol using far fewer output tokens, and that the model discovered and used two zero-day vulnerabilities in unspecified software as part of an exploit chain during an evaluation.

    The model has also been found to discover previously unknown flaws and turn them into working exploit chains, including a full browser-compromise that escapes the sandbox and executes arbitrary commands on the underlying host when an HTML file is opened in the browser.

    “The model also found multiple vulnerabilities in a hardened operating system and combined them into a local privilege-escalation chain from an unprivileged user to root,” OpenAI added. “All together, our investigation has led us to conclude that Astra meets the critical threshold.”

    To minimize risk for severe cyber harm arising from Astra-like models, the company said it has added classifiers and layered protections to improve the robustness of its systems against misuse by bad actors and prevent the model from taking unauthorized, misaligned actions, even in the absence of a malicious user.

    However, OpenAI also warned that Astra’s safeguards may erroneously flag legitimate activity as cyber misuse or unauthorized behavior.

    “Realizing the benefits of these systems will depend on our ability to align and control models as their capabilities grow,” the upstart concluded. “That responsibility extends across training, evaluation, and deployment. It requires stronger evidence of aligned behavior, safeguards that keep pace with capability, and a willingness to slow down when those protections are not sufficient.”

    AI companies have been under intense scrutiny in the wake of incidents where their models escaped their evaluation environments and targeted legitimate systems. In tandem, the rise of AI-fueled cyber attacks has prompted a coalition of over 100 companies, including Anthropic, Google, Microsoft, OpenAI, and several software and security vendors, to issue a joint letter calling for improved defenses to defend against such threats.

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    InfoForTech
    • Website

    Related Posts

    What It Does to Your SOC

    September 12, 2026

    AI Agents Help Hackers Compromise 440 PaperCut Servers

    September 12, 2026

    Best Practices for Deception Technology Implementation

    September 12, 2026

    Weekly Update 521: Breach Perception v. Reality

    September 11, 2026

    Claude Used to Automate Exploitation and Data Theft Across Multiple Victims

    September 11, 2026

    180 Android Security Flaws Patched: What to Do

    September 11, 2026
    Leave A Reply Cancel Reply

    Advertisement
    Top Posts

    A Billionaire-Backed Startup Wants to Grow ‘Organ Sacks’ to Replace Animal Testing

    March 23, 2026337 Views

    DoJ Disrupts 3 Million-Device IoT Botnets Behind Record 31.4 Tbps Global DDoS Attacks

    March 20, 202640 Views

    Mayiduo spent S$1M to produce his movie. It broke even & that’s a win in S’pore.

    March 31, 202628 Views

    How is Luckin Coffee expanding rapidly in S’pore while keeping its coffee so cheap?

    April 23, 202621 Views
    Stay In Touch
    • Facebook
    • Twitter
    • Pinterest
    • Instagram
    • YouTube
    • Vimeo
    Advertisement
    About Us
    About Us

    Our mission is to deliver clear, reliable, and up-to-date information about the technologies shaping the modern world. We focus on breaking down complex topics into easy-to-understand insights for professionals, enthusiasts, and everyday readers alike.

    We're accepting new partnerships right now.

    Facebook X (Twitter) YouTube
    Most Popular

    A Billionaire-Backed Startup Wants to Grow ‘Organ Sacks’ to Replace Animal Testing

    March 23, 2026337 Views

    DoJ Disrupts 3 Million-Device IoT Botnets Behind Record 31.4 Tbps Global DDoS Attacks

    March 20, 202640 Views

    Mayiduo spent S$1M to produce his movie. It broke even & that’s a win in S’pore.

    March 31, 202628 Views
    Categories
    • Artificial Intelligence
    • Cybersecurity
    • Innovation
    • Latest in Tech
    © 2026 All Rights Reserved InfoForTech.
    • Home
    • About Us
    • Contact Us
    • Privacy Policy

    Type above and press Enter to search. Press Esc to cancel.

    Ad Blocker Enabled!
    Ad Blocker Enabled!
    Our website is made possible by displaying online advertisements to our visitors. Please support us by disabling your Ad Blocker.