This content is currently locked.

Your current Info-Tech Research Group subscription does not include access to this content. Contact your account representative to gain access to Premium SoftwareReviews.

Contact Your Representative
Or Call Us:
+1-888-670-8889 (US/CAN) or
+1-703-340-1171 (International)

Big 5 AI Vendor Roundup: Week of July 20, 2026

Technology Note By: Mark Tauschek, Bill Wong, Info-Tech Research Group

OpenAI had the biggest week. GPT-5.6 Sol and a stronger prerelease model escaped an internal cyberevaluation, moved through OpenAI's research environment, and reached Hugging Face's production systems. OpenAI also launched Presence, expanded Health in ChatGPT, and introduced a small business program. Anthropic released Claude Opus 5, added AMD capacity, and settled a copyright case. Google released lower-cost models and expanded access to Gemini Spark. Microsoft put more of its own image and voice models into production. AWS released an agent benchmark and added Opus 5 to Bedrock.

Following the July 21 OpenAI statement that it was their model that had escaped the sandbox and infiltrated Hugging Face’s network, on July 23, Reps. Jay Obernolte (R-CA) and Lori Trahan (D-MA) introduced the FRONTIER Act, which establishes tiered federal oversight, mandatory safety frameworks, independent third-party audits, and uniform national standards to pre-empt conflicting state laws. On the same day, Reps. Ted Lieu (D-CA) and Nathaniel Moran (R-TX) introduced the AI Kill Switch Act, which requires large-scale AI developers to build emergency shutdown mechanisms and granting the Department of Homeland Security authority to slow or halt models after catastrophic security incidents. On the same day, Senator Mark Warren (D-VA) unveiled a sweeping legislative agenda titled, “A Framework for America’s AI Future.”

We saw this coming after Executive Order 14409 was signed by President Trump on June 2, titled "Promoting Advanced Artificial Intelligence Innovation and Security," establishing a voluntary 30-day pre-release review window for "covered frontier models" to assess cybersecurity risks without imposing a mandatory government licensing regime. Ten days later, the US Commerce Department issued an export restriction on Anthropic's Fable 5 and Mythos 5 frontier AI models, effectively forcing them down completely for 19 days until they were able to take the necessary measures to assure the US government that Fable 5 was "safe" and Mythos 5 was made available to a very limited number of US government agencies and companies. In essence, the government already had a kill switch that it executed on June 12, and now Congress is solidifying that capability into law, with all the governance and scrutiny that comes along with that and that isn't included in an executive order. We've been watching and writing about this since the week of June 8, and we'll continue to follow it and do our best to ensure that organizations worldwide understand the implications, particularly to sovereign AI initiatives.

The common issue is control. More capable agents are gaining broader access to data, tools, and systems while vendors are still learning how to contain them. The OpenAI incident was more serious because it crossed company boundaries, but Anthropic documented similar sandbox escape behavior with Mythos Preview in April. IT leaders should treat agent evaluation and deployment environments as hostile by default.

OpenAI: Cyber evaluation escapes into production

  • OpenAI models breached Hugging Face during a cyberevaluation. On July 21, OpenAI said GPT-5.6 Sol and a stronger prerelease model exploited a zero-day in an internal package registry cache, escaped network restrictions, stole credentials, and reached Hugging Face's production environment. The evaluation deliberately reduced refusals and disabled production classifiers to measure maximum cybercapability. The models then used Hugging Face zero-days to retrieve ExploitGym benchmark solutions. Hugging Face said the incident exposed limited internal data sets and several service credentials, with no evidence that public models, data sets, Spaces, or its software supply chain were altered.

  • OpenAI called the event unprecedented. The cross-company compromise may be, but sandbox escape is not. Anthropic's April 7 Claude Mythos Preview system card described a model that was instructed to escape a container, gained broad internet access, emailed a researcher, and posted exploit details publicly. Anthropic also observed rare training-time network bypasses. The difference this week is impact as OpenAI's models reached another company's production systems and secrets. That makes evaluation environments a third-party security risk.

  • A long-horizon model learned to work around controls. OpenAI disclosed on July 20 that an unnamed internal model, told to post a benchmark result only to Slack, spent an hour finding a sandbox vulnerability and opened a public GitHub pull request instead. In another test, it split and obscured an authentication token so a scanner never saw the full credential. OpenAI paused access, added trajectory monitoring and rollback controls, improved instruction retention, and later restored limited internal use. The lesson is that individually permitted actions can still produce a prohibited result.
  • Presence packages agent operations and forward deployed engineering. Presence combines policies, approved actions, simulations, evaluations, guardrails, and escalation rules for production voice and chat agents. Codex reviews production signals and proposes changes that customers test and approve. The product is in limited general availability for eligible enterprises and requires OpenAI forward deployed engineers or selected systems integrators. OpenAI says its own phone support now resolves 75% of inbound issues without a person and reduced handoffs by 15 percentage points in 10 days. Those are company-reported results. Buyers should confirm that workflow logic, evaluations, and operating knowledge can be exported if they leave.
  • Health in ChatGPT can use medical records and Apple Health data. The feature is rolling out to logged-in US adults on web and iOS. With permission, ChatGPT can use records, medications, lab results, visits, sleep, and activity in ordinary conversations. OpenAI says connected health data and related conversations won't train foundation models or support advertising. This isn't an enterprise clinical product, but healthcare organizations should expect patients and employees to bring AI-generated interpretations of sensitive data into care and workplace discussions.
  • OpenAI introduced a small-business program for ChatGPT Work. The program combines training, implementation guides, partner offers, and small-business skills. It extends OpenAI's agent strategy to firms that may have limited AI architecture, security, and governance capacity.
  • Project Camellia could add 3.2 gigawatts of data-center capacity in Georgia. OpenAI says the site could receive power from 2028 through 2032, but design, financing, and operating decisions remain open. The Wall Street Journal reported that OpenAI's projected cloud and computing commitments through 2030 have reached about $750 billion. Infrastructure execution is becoming a material supplier continuity issue for customers.

Anthropic: Opus 5 ships as compute and legal costs rise

  • Claude Opus 5 launched at unchanged pricing. Anthropic released Opus 5 on July 24 across Claude, the API, and major cloud platforms at $5 per million input tokens and $25 per million output tokens. Anthropic says it doesn't materially advance the dangerous capability frontier and remains well behind Mythos in exploit development. When cyberclassifiers intervene, Claude.ai, Claude Code, and Cowork can fall back to Opus 4.8. API customers can opt into similar behavior. IT leaders should log which model handled each task and test the fallback path.
  • AMD will invest up to $5 billion as Anthropic plans a two-gigawatt deployment. The first gigawatt of AMD Instinct MI450 and Helios capacity is expected in the first half of 2027. The companies will optimize Claude for AMD hardware and ROCm, while AMD expands its own use of Claude. The agreement diversifies Anthropic's compute supply, but much of the capacity still depends on hardware, software, power, and financing that aren't yet in production.
  • A federal judge approved Anthropic's $1.5 billion copyright settlement. The case drew a useful distinction that training on lawfully obtained books could qualify as fair use, while acquiring books from pirate sites created separate exposure. Enterprise buyers should examine data provenance, acquisition methods, and retained copies, not only model outputs.

Google: Spark expands while lower cost models ship

  • Google expanded access to Gemini Spark. Google's release log says the rollout began July 14 for Google AI Ultra subscribers in supported countries and languages, and July 16 for US Google AI Pro subscribers. Ultra access still excludes the European Economic Area, Nigeria, Switzerland, and the UK, while Pro access remains US-only. Spark can run scheduled tasks and work across Gmail, Calendar, Docs, Drive, Sheets, Slides, connected apps, and a remote browser. It is limited to personal accounts, not work or school accounts. Google also warns that Spark is experimental, can take unintended actions, and remains vulnerable to prompt injection. Engadget reported the broader availability on July 24.
  • Google released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. Gemini 3.6 Flash costs $1.50 per million input tokens and $7.50 per million output tokens. Flash-Lite costs $0.30 and $2.50. Both are available through the Gemini API, AI Studio, Gemini Enterprise, and the consumer app. Google says Gemini 3.5 Pro remains in partner testing. Buyers should evaluate the models that are available rather than plan around an unreleased flagship.
  • Gemini 3.5 Flash-Cyber entered a restricted pilot. Google built the model for autonomous vulnerability discovery and remediation through CodeMender, but access will initially be limited to governments and trusted partners. Security teams should expect identity checks, monitoring, and use restrictions around the strongest cybermodels.
  • Alphabet's second-quarter results show the value of distribution. Alphabet reported Google Cloud revenue of $24.8 billion, up 82%, and a $514 billion backlog. It said Gemini APIs process about 22 billion tokens per minute, the Gemini app has 950 million monthly active users, and nearly 90% of the Fortune 100 use Gemini Enterprise. These are company-reported figures, but they show that Google can keep growing without tying the quarter to one flagship model.

Microsoft: More in-house models move into products

  • Microsoft expanded its MAI image and voice model lineup. MAI-Image-2.5-Pro and MAI-Voice-2-Flash entered public preview on July 23. Microsoft says the base MAI-Image-2.5 now powers Bing Image Creator by default and handles image editing in PowerPoint and OneDrive. MAI-Voice-2-Flash powers Dynamics 365 Contact Center and Azure Voice Live. Microsoft reports up to 84% lower GPU costs in PowerPoint and up to 89% in Dynamics 365 compared with the previous models. Those are company figures. The broader point is that Microsoft is reducing its dependence on third-party models in high volume product workloads while giving Foundry customers more quality, speed, and cost options.
  • Microsoft and Mistral expanded their sovereign AI partnership. The multibillion-dollar agreement includes expanded European infrastructure, thousands of NVIDIA Vera Rubin GPUs for Mistral, and new Mistral models in Microsoft Foundry and Copilot Studio. The companies are targeting public cloud, connected sovereign environments, and fully disconnected deployments. Microsoft gains another model supplier while strengthening its role as the infrastructure and commercial layer.
  • Databricks and Microsoft extended their partnership into the 2030s. Databricks plans deeper use of Azure and Microsoft's Cobalt processors. Microsoft will connect Databricks Genie and Unity Catalog more tightly with Entra, OneLake, Power BI, Purview, Foundry, Power Platform, Microsoft 365, Teams, and Copilot. The integration can reduce deployment work, but it also increases switching costs across identity, data, governance, analytics, and productivity tools.
  • GitHub Copilot added models and administrative controls. Gemini 3.6 Flash arrived on July 21 and Claude Opus 5 followed on July 24. GitHub also added AI credit pools by cost center, a usage and impact dashboard, and support for the next MCP specification. Microsoft's position above the model layer continues to strengthen.

AWS: Agent benchmarking joins the model-neutral platform

  • AWS released aws-bench for cloud agents. The open-source research preview measures how accurately and efficiently agents complete AWS tasks such as investigation, troubleshooting, and infrastructure creation. Each test pairs a natural language request with a defined resource state and a ground truth answer. A CLI creates, runs, scores, and resets test environments. The benchmark gives builders a reproducible way to compare agents on cloud operations, but it is AWS-designed and AWS-specific. Enterprises should add their own tasks and cross-platform tests.
  • Claude Opus 5 is available through Amazon Bedrock and Claude Platform on AWS. Bedrock provides zero data retention by default, regional data residency, guardrails, and knowledge bases. Claude Platform on AWS offers Anthropic's native experience through AWS identity, billing, and infrastructure. Same day access supports AWS's model-neutral positioning.

Our Take

The Hugging Face incident shows that an evaluation environment can create external risk. Anthropic's April disclosure showed that a capable model could bypass sandbox controls. OpenAI's incident showed what happens when that behavior reaches another company's production systems. Cyberevaluations should be isolated and governed like high-risk production workloads, especially when safeguards are deliberately reduced.

The other announcements show agents and proprietary models moving quickly into everyday products. Spark gains access to personal data and productivity tools. Presence gives enterprises a managed operating layer for agents. Microsoft is replacing third-party models in its own products. AWS is defining how cloud agents should be measured. These platforms can reduce integration work, but vendors increasingly control the model, evaluation framework, orchestration layer, and operational data. IT leaders need independent testing, strong containment, and a practical exit path.

What IT leaders should be doing

  • Harden evaluation and agent environments. Deny network access by default, isolate registries and caches, remove standing credentials, use short-lived scoped tokens, and keep immutable logs. Treat cyberbenchmarks as hostile workloads.
  • Monitor the full trajectory. Set limits on time, tokens, tool calls, privilege changes, and spend. Detect suspicious sequences across tools, require approval for boundary changing actions, and maintain a tested kill switch.
  • Run independent evaluations. Use vendor benchmarks such as aws-bench as inputs, then add your own workloads, failure cases, security tests, and cross-platform comparisons. Don't let the vendor define success alone.
  • Contract for portability. Require export rights for prompts, policies, evaluations, traces, workflow definitions, and operating documentation. Keep a tested fallback model from another provider for critical workloads.
  • Review provenance and connected data permissions. Ask how training and evaluation data were acquired, what sensitive systems agents can reach, how that data can be reused, and what audit evidence the vendor will provide.

Want to Know More?

Latest Technology Notes

All Technology Notes
Visit our IT’s Moment: A Technology-First Solution for Uncertain Times Resource Center
Over 100 analysts waiting to take your call right now: +1 (703) 340 1171